INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     skill
    -0.07
    PLIER
    -0.07
    -0.06
    utschein
    -0.06
     phái
    -0.06
     terminology
    -0.06
     теперь
    -0.06
     drawing
    -0.06
     nums
    -0.06
     Notification
    -0.06
    POSITIVE LOGITS
    wanted
    0.08
     fueled
    0.07
    ahrung
    0.06
    kat
    0.06
    andoned
    0.06
     عملکرد
    0.06
    Ans
    0.06
     FLASH
    0.06
    HY
    0.06
    Reject
    0.06
    Act Density 0.005%

    No Known Activations