INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     Někter
    -0.07
     Nurse
    -0.07
    observ
    -0.07
     morphology
    -0.07
     بتن
    -0.07
     Meredith
    -0.07
     maths
    -0.06
     Aph
    -0.06
     Hannah
    -0.06
     Herb
    -0.06
    POSITIVE LOGITS
     exciting
    0.13
     excitement
    0.12
     excited
    0.10
     Exc
    0.08
    ovic
    0.08
     velocity
    0.08
     Tickets
    0.08
    Action
    0.07
    DX
    0.07
    Tickets
    0.07
    Act Density 0.015%

    No Known Activations