INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    inux
    -0.15
    ingham
    -0.14
     Poz
    -0.14
    atk
    -0.14
    ики
    -0.14
    agne
    -0.14
    LowerCase
    -0.13
    лаж
    -0.13
    symbols
    -0.13
    rink
    -0.13
    POSITIVE LOGITS
    erah
    0.17
    ercul
    0.17
     Fus
    0.16
    mium
    0.15
    aras
    0.14
    ç±
    0.14
    -REAL
    0.14
    cona
    0.14
    aland
    0.14
     bif
    0.14
    Act Density 1.917%

    No Known Activations