INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    0.40
     περισσότε
    0.39
     powerAll
    0.38
     automático
    0.38
     kinda
    0.37
    وروب
    0.36
     ಕೇಂದ್ರ
    0.36
     próxim
    0.36
    तावनी
    0.36
     tất
    0.35
    POSITIVE LOGITS
    perms
    0.37
     n
    0.37
    perm
    0.36
    ro
    0.36
    krit
    0.35
    verbal
    0.35
    java
    0.35
    ysel
    0.34
                
    0.34
    ѝ
    0.34
    Act Density 0.008%

    No Known Activations