INDEX
    Explanations
    New Auto-Interp
    Negative Logits
     bekl
    -0.17
     komp
    -0.15
    ustr
    -0.15
    ppe
    -0.14
    ebin
    -0.14
    æľ
    -0.14
    isel
    -0.14
     kapit
    -0.14
     lẽ
    -0.13
    edin
    -0.13
    POSITIVE LOGITS
    ãĥ¼ãĥĢ
    0.16
    ces
    0.15
    vic
    0.15
    cla
    0.14
    vk
    0.14
     grunt
    0.14
    inden
    0.14
    wang
    0.14
    Ùĥز
    0.14
    ine
    0.14
    Act Density 0.012%

    No Known Activations