INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    chair
    -0.07
    arrow
    -0.07
     Fay
    -0.07
    -0.07
     Kale
    -0.07
     Sok
    -0.07
    -0.07
    โช
    -0.07
     pym
    -0.06
     Roh
    -0.06
    POSITIVE LOGITS
     must
    0.13
    must
    0.10
     Must
    0.09
    Must
    0.08
     MUST
    0.08
    必须
    0.07
     should
    0.07
    .tests
    0.07
    ebiliriz
    0.07
    =models
    0.07
    Act Density 0.049%

    No Known Activations