INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    /Table
    -0.17
     enc
    -0.15
    ertz
    -0.14
     disin
    -0.14
    ama
    -0.14
    Disposable
    -0.13
     pun
    -0.13
    bero
    -0.13
    -O
    -0.13
    /Edit
    -0.13
    POSITIVE LOGITS
    ibble
    0.14
    VN
    0.14
    moon
    0.14
    šak
    0.14
    Äįet
    0.14
    ÅŁk
    0.14
    orca
    0.14
    urance
    0.14
    irty
    0.14
    nil
    0.13
    Act Density 0.040%

    No Known Activations