INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    -0.08
     받을
    -0.08
     almak
    -0.08
    -0.08
    -0.08
     Selbst
    -0.08
    sorting
    -0.08
     六合
    -0.08
    sorted
    -0.07
    自主
    -0.07
    POSITIVE LOGITS
     idi
    0.11
    0.09
    .Exceptions
    0.09
     sayings
    0.08
     collo
    0.08
     slang
    0.08
     Idi
    0.08
     contractions
    0.08
    Idi
    0.08
     lifestyle
    0.08
    Act Density 0.011%

    No Known Activations