INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     Worship
    -0.07
     olduğu
    -0.07
     worship
    -0.07
     тур
    -0.06
     chemistry
    -0.06
     fixing
    -0.06
     pragmatic
    -0.06
    ERTICAL
    -0.06
    çe
    -0.06
     Phone
    -0.06
    POSITIVE LOGITS
     màu
    0.07
    高速
    0.07
     opi
    0.06
    ="(
    0.06
    "]/
    0.06
    0.06
     angi
    0.06
    وئ
    0.06
    _DEL
    0.06
     strang
    0.06
    Act Density 0.011%

    No Known Activations