INDEX
    Explanations
    New Auto-Interp
    Negative Logits
    ores
    -0.19
    erator
    -0.17
    tings
    -0.17
    ework
    -0.17
    cis
    -0.16
    tf
    -0.15
    ouncements
    -0.15
    ÑĤаÑĢ
    -0.15
    еÑĢг
    -0.15
    czy
    -0.15
    POSITIVE LOGITS
     Illustrated
    0.24
    manship
    0.21
    book
    0.20
    books
    0.19
     medicine
    0.19
    medicine
    0.19
     illustrated
    0.19
     Medicine
    0.19
    net
    0.18
     betting
    0.18
    Act Density 0.014%

    No Known Activations