INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    (dictionary
    -0.07
    github
    -0.07
    leyen
    -0.07
    ladık
    -0.06
    BaseUrl
    -0.06
    ayım
    -0.06
    Booking
    -0.06
    W
    -0.06
    ((_
    -0.06
     ------------
    -0.06
    POSITIVE LOGITS
    FIN
    0.08
    plementary
    0.07
    	Global
    0.07
    0.07
     shark
    0.07
     '')↵
    0.06
    ithmetic
    0.06
     Analy
    0.06
    )];
    ↵
    0.06
    0.06
    Act Density 0.006%

    No Known Activations