INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     cor
    -0.07
    Component
    -0.06
    Agent
    -0.06
     fuel
    -0.06
     diamond
    -0.06
     bears
    -0.06
    (account
    -0.06
     ti
    -0.06
     FCC
    -0.06
    นก
    -0.06
    POSITIVE LOGITS
    okus
    0.07
    ы
    0.06
    akter
    0.06
    інь
    0.06
    ало
    0.06
    αλ
    0.06
    Power
    0.06
     αυτή
    0.06
     Fixes
    0.06
    instead
    0.06
    Act Density 0.003%

    No Known Activations