INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     Zodiac
    -0.08
    ambling
    -0.08
     자세
    -0.07
    рап
    -0.07
     подачи
    -0.07
     school's
    -0.07
    smanship
    -0.07
     raccord
    -0.07
    -0.07
     Cobra
    -0.07
    POSITIVE LOGITS
    _LAYER
    0.08
     ফুল
    0.08
     Layer
    0.08
    0.07
     Melbourne
    0.07
     fem
    0.07
     Provence
    0.07
     MIS
    0.07
    üyor
    0.07
    pap
    0.07
    Act Density 0.001%

    No Known Activations