INDEX
    Explanations

    Sentence beginnings

    New Auto-Interp
    Negative Logits
    -0.07
    Maximum
    -0.07
    Radius
    -0.06
    _no
    -0.06
     indications
    -0.06
    Status
    -0.06
     Netanyahu
    -0.06
     Lease
    -0.06
     dinh
    -0.06
     cocci
    -0.06
    POSITIVE LOGITS
     neden
    0.07
    fell
    0.06
     заключ
    0.06
    ockey
    0.06
    .visitMethod
    0.06
    INF
    0.06
     його
    0.06
    chluss
    0.06
    ]<<
    0.06
    0.06
    Act Density 0.046%

    No Known Activations