Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting
35 findings
Linear Probing / Ridge regression linear probing / Linear probe / Linear probe fine-tuning / Linear regression probing / Linear ridge regression probes / Supervised probing / ERM linear probe
23 findings
Principal component analysis
21 findings
ROME
18 findings
Sparse autoencoder / Sparse autoencoders / K-sparse autoencoder / Topk sparse autoencoder / Scaling and Evaluating Sparse Autoencoders / Cunningham et al. 2023 (sparse autoencoders)
17 findings
Activation patching / Activation replacement / Cross-model activation patching (CMap)
16 findings
GCG
15 findings
In-Context Learning / In-context learning prompt
14 findings
Logit lens
14 findings
LoRA
13 findings
MEMIT
13 findings
ReAct
13 findings
Self-Consistency / Self-consistency prompting / Wang et al. 2023 (self-consistency) / Wang et al. 2023b (Self-Consistency)
13 findings
Cosine similarity / Cosine similarity analysis / Cosine semantic similarity / cosine similarity of hidden states / Sample-wise cosine similarity / Cosine similarity of attention maps / Cosine similarity perturbation analysis / Cosine similarity template matching / Cosine similarity to neighbours / Semantic consistency (cosine similarity)
12 findings
Expected Calibration Error / Integral Calibration Error (ECE)
12 findings
Integrated Gradients / Integral of gradients
10 findings
Mean Ablation / direct effect mean ablation / Mean token ablation
10 findings
Probing classifiers / MLP probing classifiers / Q16 classifier
10 findings
Activation steering / Mean steering / PCA steering
9 findings
CLIPScore
9 findings
Causal mediation analysis / Vig et al. 2020 (causal mediation analysis)
9 findings
Path Patching
9 findings
Zero-shot Chain-of-Thought / Wei et al. 2022 (Chain of Thought) / Zero-shot chain-of-thought prompting
9 findings
Direct Preference Optimization / Direct Preference Optimisation / DPO / Rafailov et al. 2023 (Direct Preference Optimization)
8 findings
PAIR
8 findings
Few-shot prompting / 2-shot prompting / Few-shot ICL / Few-shot prompting for base models
7 findings
Grad-CAM
7 findings
Self-Refine
7 findings
SparseGPT
7 findings
Wanda
7 findings
Zero-shot prompting
7 findings
BLEU / BLEU@4
6 findings
FID
6 findings
Logistic regression / Linear decoder (logistic regression) / Logistic classification / Logistic regression head
6 findings
MEND
6 findings
PPO
6 findings
Retrieval-Augmented Generation / Retrieval augmentation (top-5 chunks)
6 findings
Spearman rank correlation
6 findings
AutoDan
5 findings
CHAIR / CHAIRi / CHAIRS
5 findings
Cohen's kappa
5 findings
ICE
5 findings
Inference-Time Intervention / ITI / Neuron intervention (pinning activation)
5 findings
ROUGE-L
5 findings
Representation Engineering / Representation engineering (control vectors) / Zou et al. 2023 (Representation Engineering) / Zou et al. (representation engineering)
5 findings
k-nearest neighbours classifier / Nearest-neighbor baseline
5 findings
pass@k / pass n@k / pass@1 / pass@t
5 findings
t-SNE
5 findings
ABC
4 findings
AUROC
4 findings
CIDER (Compactness and Dispersion for OOD detection)
4 findings
Causal Tracing
4 findings
Conceptor
4 findings
Distributed Alignment Search / DAS
4 findings
DoLa
4 findings
EK-FAC Influence Functions
4 findings
GRACE
4 findings
GeoEx feature-plane extraction
4 findings
Hypothesis Refinement / Iterative hypothesis refinement
4 findings
K-means clustering
4 findings
LLM-as-a-Judge / GPT-4 as judge / GPT-4o as LLM judge
4 findings
LM Evaluation Harness
4 findings
Layer-Selective Rank Reduction / LASER
4 findings
Least-to-Most Prompting
4 findings
PGD (Projected Gradient Descent)
4 findings
Pearson correlation
4 findings
Perspective API
4 findings
Pointing game
4 findings
Representational Similarity Analysis
4 findings
SOTOPIA-EVAL
4 findings
Verbalized Confidence / Verbal confidence elicitation
4 findings
Arithmetic Coding
3 findings
BERTScore
3 findings
Benjamini-Hochberg FDR correction
3 findings
Brier score
3 findings
CAUSE
3 findings
CHRF
3 findings
CKA
3 findings
CLIP
3 findings
COMET
3 findings
Calibrate Before Use / Zhao et al. 2021 (Calibrate Before Use)
3 findings
Contrastive Pairs Method
3 findings
Fast-DetectGPT
3 findings
GPTQ
3 findings
ICD
3 findings
Linear CKA
3 findings
MM-SHAP
3 findings
Magnitude Pruning
3 findings
Multilayer perceptron
3 findings
Nucleus Sampling
3 findings
OpenCLIP
3 findings
Residual Stream Patching
3 findings
Retrobust
3 findings
Ridge regression / Bootstrap ridge regression
3 findings
Spectral Clustering
3 findings
System-embedded diffusion bridges
3 findings
TextSpan
3 findings
TruthX
3 findings
Tuned Lens
3 findings
Uniform TTM
3 findings
Weighted Banzhaf interaction index
3 findings
Weighted Cohen's Kappa
3 findings
XGBoost
3 findings
gzip
3 findings
ASH
2 findings
Arena-Hard-Auto
2 findings
AutoAttack
2 findings
Behavior Cloning
2 findings
BenchCLAMP
2 findings
CRITIC
2 findings
Classifier-free guidance / Classifier-free diffusion guidance
2 findings
Constrained Decoding
2 findings
Control task with shuffled labels
2 findings
DECAF
2 findings
DeiT
2 findings
E5-Mistral
2 findings
EPIC distance
2 findings
ESMFold
2 findings
F1 score
2 findings
FAISS
2 findings
FGSM (Fast Gradient Sign Method)
2 findings
FLIPD
2 findings
G-Eval
2 findings
Gemma Scope
2 findings
HumanJailbreaks
2 findings
Interchange Intervention
2 findings
KN Edit
2 findings
Kendall's tau / Kendall-tau
2 findings
Knowledge Attribution
2 findings
Knowledge Neurons / Dai et al. 2022 (Knowledge Neurons)
2 findings
Köppen-Geiger climate classification
2 findings
Linear Relational Embedding
2 findings
Linear regression
2 findings
Llama-Guard-3-8B
2 findings
MI-Zero
2 findings
MLM Scoring
2 findings
MMseqs2
2 findings
MPPi
2 findings
Mean-difference concept vector
2 findings
MoleculeSTM
2 findings
Moran's I
2 findings
Multidimensional Scaling
2 findings
NNSight and NDIF
2 findings
Norm-based Analysis
2 findings
NudeNet
2 findings
OTI
2 findings
OpenAI Moderation API
2 findings
OrdinalCLIP
2 findings
Orthogonal Matching Pursuit
2 findings
PEGASUS
2 findings
PVQ-RR
2 findings
Paired t-test
2 findings
Prefix Tuning / Li & Liang 2021 (Prefix-tuning)
2 findings
ProteinCLAP-EBM-NCE
2 findings
Q-learning
2 findings
RAG
2 findings
RAHF (Rich Human Feedback) score
2 findings
Random Smoothing
2 findings
SAS Probe
2 findings
SSCD
2 findings
SSPAttack
2 findings
Self-RAG
2 findings
Sequential Monte Carlo
2 findings
Shapiro-Wilk test
2 findings
StreamingLLM
2 findings
TAP (Tree of Attacks with Pruning)
2 findings
Temperature Scaling
2 findings
UMAP
2 findings
UTMOS
2 findings
VAL
2 findings
Vendi score
2 findings
Visual Prompt Tuning
2 findings
n-Shapley values
2 findings
ACDC (Automated Circuit Discovery)
1 finding
ADRound
1 finding
APE
1 finding
AWQ
1 finding
Activation Addition / Turner et al. (activation addition)
1 finding
Adversarial Gradient Integration
1 finding
Agglomerative Clustering
1 finding
Area between insertion and deletion curves
1 finding
AttackVLM
1 finding
Attention knockout
1 finding
Attribution Patching
1 finding
AugMix
1 finding
AutoGPT
1 finding
Autointerpretability
1 finding
BAP
1 finding
BGE-large-en-v1.5
1 finding
BLEURT
1 finding
BM25 / BM25 retrieval
1 finding
BRECQ
1 finding
Batch Calibration
1 finding
Best-of-N Sampling
1 finding
C&W
1 finding
CATS
1 finding
CC-SHAP
1 finding
CLIP-Dissect
1 finding
CLIPSelf
1 finding
CRAG
1 finding
Centered Kernel Alignment
1 finding
ChatCaptioner
1 finding
CipherChat
1 finding
Circular correlation
1 finding
CoOp
1 finding
CodeBERTScore
1 finding
CodeBLEU
1 finding
Contextual Calibration
1 finding
Continual Knowledge Learning
1 finding
Corrupting CoT
1 finding
CutMix
1 finding
DBNet
1 finding
DFQ
1 finding
DIPPER-paraphraser-xxl
1 finding
DQN
1 finding
DRTune
1 finding
Decomposed Prompting
1 finding
DeepFloyd IF
1 finding
DeepFool
1 finding
Detoxify
1 finding
Dijkstra shortest path
1 finding
EATA
1 finding
EBO
1 finding
EasyOCR
1 finding
Empirical Neural Tangent Kernel
1 finding
Equal Opportunity Difference
1 finding
Error Consistency
1 finding
FLAC
1 finding
FairFace
1 finding
Faith-Shap
1 finding
Feature Visualization by Optimization
1 finding
Fleiss' kappa / Fleiss-κ
1 finding
Foot-in-the-Door Attack / Foot-in-the-door (Wang et al. 2024)
1 finding
Frozen in Time
1 finding
GEDI
1 finding
GLIP
1 finding
GPT-4-32k
1 finding
GRIDE
1 finding
GWL Test
1 finding
Gaussian KDE
1 finding
Getis-Ord Gi*
1 finding
Gini Index
1 finding
Giphy Celebrity Detector
1 finding
Grad-ECLIP
1 finding
Greedy Maximum Coverage Approximation
1 finding
GroundingDINO
1 finding
Gzip Compression
1 finding
H2O
1 finding
H3 hexagonal grid
1 finding
HSIC
1 finding
Head masking
1 finding
Hybrid geo-latent graph
1 finding
HyperDAS
1 finding
I/O CoT
1 finding
IGWL test
1 finding
INPCA
1 finding
ImgJP
1 finding
In-Context Knowledge Editing / IKE
1 finding
Influence Functions / Grosse et al. 2023 (influence functions)
1 finding
Info-RAG
1 finding
Information Flow Routes / Ferrando & Voita 2024 (information flow routes)
1 finding
Information Imbalance
1 finding
InstructZero
1 finding
Isotonic regression
1 finding
Jaccard similarity
1 finding
Jensen-Shannon divergence
1 finding
KID
1 finding
KL Divergence
1 finding
Kendall's τ
1 finding
Knowledge Distillation
1 finding
Knowledge Editor / KE
1 finding
Kolmogorov-Smirnov two-sample test
1 finding
Kronfluence
1 finding
L1 Magnitude Pruning
1 finding
LLM-Pruner
1 finding
LLM.int8()
1 finding
LLaVA-Next
1 finding
LOSt
1 finding
LPIPS
1 finding
LWP
1 finding
Lasso regression
1 finding
Ledoit-Wolf Shrinkage Estimator
1 finding
Levenshtein distance
1 finding
Linear Assignment Problem
1 finding
Linear HSIC
1 finding
Linear support vector machine
1 finding
Llama 3
1 finding
Local Moran's I (LISA)
1 finding
Local fidelity R2 over the feature lattice
1 finding
Logit Anchoring
1 finding
M2IB
1 finding
MAUVE
1 finding
MDS (Mahalanobis Distance-based Score)
1 finding
MEMO
1 finding
METEOR
1 finding
MHCFlurry 2.0
1 finding
MMP
1 finding
MPNet
1 finding
MSP (Maximum Softmax Probability)
1 finding
McNemar's test
1 finding
MiniLM-L6-V2
1 finding
Multi-Agent Debate
1 finding
Multi-task Distributed Alignment Search / MDAS
1 finding
Mutual Nearest-Neighbor Kernel Alignment
1 finding
NLPO
1 finding
NNSight
1 finding
NNSight/NDIF
1 finding
NUPES
1 finding
Net2Brain
1 finding
Network Dissection
1 finding
Normalized Kendall-tau distance / Normalized Kendall-τ distance
1 finding
NudNet
1 finding
Null-text inversion
1 finding
ODIN
1 finding
OPTQ
1 finding
Onion
1 finding
Online DPO
1 finding
OpenMask3D
1 finding
PAC-Bayes bound
1 finding
PAC-Bayesian Analysis
1 finding
PAL (Program-Aided Language Models) / PaL
1 finding
PAP
1 finding
PSNR
1 finding
Partial-LRP
1 finding
Path-norm
1 finding
Pearson correlation coefficient
1 finding
PerSAM
1 finding
Permutation test
1 finding
Persona Modulation
1 finding
Platt scaling
1 finding
Poison-RLHF
1 finding
Polynomial Regression
1 finding
PowerQuant
1 finding
Procrustes Analysis
1 finding
Program of Thoughts
1 finding
Prompt Ensembling
1 finding
Prompt Tuning / Soft prompt tuning
1 finding
QF-Attack
1 finding
RALF
1 finding
RED++
1 finding
REX
1 finding
RLOO
1 finding
ROAR
1 finding
ROUGE-1
1 finding
Random Forest Regression
1 finding
Rays
1 finding
ReNeLLM
1 finding
Reasoning Score
1 finding
RegionCLIP
1 finding
Relative Representations
1 finding
Retriever-Dictionary module / Retriever-Dictionary (RD) module / RD module
1 finding
SERAC
1 finding
SHNAP
1 finding
SNAP
1 finding
SORT
1 finding
SPADE (Sample-efficient Probabilistic Detection)
1 finding
SQuant
1 finding
SSIM
1 finding
SVCCA
1 finding
Safely Partial-Parameter Fine-Tuning / SPPFT (Safely Partial Parameter Fine-Tuning)
1 finding
ScreenOT
1 finding
Segment Anything
1 finding
Self-BLEU
1 finding
Self-Critique
1 finding
Sentence-BERT
1 finding
Sequential Integrated Gradients
1 finding
Shapley values / SHAP
1 finding
SignFlip
1 finding
SignHunt
1 finding
Simple Gradient
1 finding
Singular Value Decomposition / Singular value decomposition of trajectory ensemble
1 finding
Smooth ECE
1 finding
SmoothLLM
1 finding
SpRAy (Spectral Relevance Analysis)
1 finding
Spearman's ρ
1 finding
Speculative Decoding
1 finding
Square Attack
1 finding
SqueezeLLM
1 finding
Structural Probing
1 finding
Structured Prompting
1 finding
Subspace activation patching
1 finding
Successor Representation / Dayan 1993 (Successor representation)
1 finding
SymbolicTOM
1 finding
TCAV (Testing with Concept Activation Vectors)
1 finding
TD-MPC
1 finding
TENT
1 finding
TIFA
1 finding
TLDR
1 finding
TextHoaxer
1 finding
Textual Inversion
1 finding
Therapeutic Antibody Profiler
1 finding
Thompson Sampling
1 finding
Tok-RAG
1 finding
TotalSegmentator
1 finding
TransformerLens
1 finding
Two-interval forced choice (2IFC)
1 finding
Two-way ANOVA with Tukey HSD
1 finding
Uniform Magnitude Pruning
1 finding
Universal Sentence Encoder
1 finding
Universal guidance
1 finding
VAA
1 finding
VIM
1 finding
Variance Ratio Test
1 finding
ViT-Slim
1 finding
Weight orthogonalization
1 finding
Wilson's confidence interval
1 finding
gSHNAP (generalized SHNAP)
1 finding
lm-eval
1 finding
Activation Projection / Probing (activation projection) / Vocabulary projection
0 findings
Architecture-Adapted Multilingual Integrated Gradients
0 findings
Attention Calibration / ACT (Attention Calibration Technique)
0 findings
Azure AI Language
0 findings
BadNets
0 findings
Best-Worst Scaling
0 findings
Binomial regression
0 findings
CCA
0 findings
CNNIQA
0 findings
Canonical Correlation Analysis
0 findings
Chi-squared test
0 findings
Circuit Breakers / Circ-Break
0 findings
Contrastive Decoding
0 findings
Counter-fitting / Counter-fitted embeddings
0 findings
Cross-Entropy Method
0 findings
CrystalBleu
0 findings
DDIM
0 findings
Deep Ensembles
0 findings
DeepZ
0 findings
Dice similarity coefficient
0 findings
Diffusion Purification
0 findings
Divergent Association Task
0 findings
Eureka
0 findings
Factor Analysis
0 findings
Fisher's r-to-z transformation
0 findings
Gaussian Smoothing
0 findings
Hamiltonian Monte Carlo
0 findings
Independent Component Analysis
0 findings
Infini-gram
0 findings
Influence Patterns
0 findings
InfoNCE
0 findings
Jaro-Winkler distance
0 findings
Knowledge Circuits
0 findings
Krippendorff's α
0 findings
LLMEval
0 findings
LLaMA 7B
0 findings
Layer-wise Relevance Propagation
0 findings
LlamaScope
0 findings
MIRo
0 findings
MPA
0 findings
Manipulate-Anything
0 findings
Multiresolution Hash Encoding
0 findings
Mutual Information
0 findings
NIMA
0 findings
OPRO
0 findings
Optimal Transport Calibration
0 findings
Parallel Context Windows / Parallel Context Window (PCW)
0 findings
ProC3S
0 findings
Proximal Policy Optimization
0 findings
QLoRA
0 findings
RAHF
0 findings
Randomized SVD
0 findings
Reward-Augmented Decoding
0 findings
Scissorhands
0 findings
SmoothGrad / Smooth Gradients
0 findings
Soft Actor-Critic
0 findings
Truth Forest
0 findings
Two-Stage Least Squares / Two-stage least squares regression
0 findings
UniPC
0 findings
VLAD
0 findings
VQAScore
0 findings
Variational Autoencoder
0 findings
WIMBD
0 findings
Whisper
0 findings
Whisper-Large-V3 / Whisper-Large-V3 ASR
0 findings
Word Mover's Distance
0 findings
ZeroCap
0 findings