INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     worship
    -0.08
    Impossible
    -0.07
    posit
    -0.07
     McCarthy
    -0.06
    LOOR
    -0.06
     Richards
    -0.06
    PMENT
    -0.06
     Deposit
    -0.06
     Hit
    -0.06
    _head
    -0.06
    POSITIVE LOGITS
    0.07
    0.07
    ουμε
    0.06
     од
    0.06
    .Design
    0.06
    came
    0.06
     externally
    0.06
    εδ
    0.06
    0.06
    ardi
    0.06
    Act Density 0.008%

    No Known Activations