INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    .Editor
    -0.08
     discusses
    -0.08
     empty
    -0.07
     compromise
    -0.07
     steal
    -0.07
     ومس
    -0.07
     stool
    -0.07
     Stockton
    -0.07
     patrol
    -0.07
     Cul
    -0.07
    POSITIVE LOGITS
    .batch
    0.10
    batch
    0.10
    =batch
    0.10
     batch
    0.10
    Batch
    0.10
    0.10
     Batch
    0.10
    Inputs
    0.09
    _inputs
    0.09
    _BATCH
    0.09
    Act Density 0.009%

    No Known Activations