- Language Model Alignment in Multilingual Trolley Problems5 findings
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models4 findings
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts4 findings
- Are Models Biased on Text without Gender-related Language?4 findings
- Benchmarking Agentic Workflow Generation4 findings
- Beyond Memorization: Violating Privacy via Inference with Large Language Models4 findings
- Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learning4 findings
- Boosting the visual interpretability of CLIP via adversarial fine-tuning4 findings
- CAS: A Probability-Based Approach for Universal Condition Alignment Score4 findings
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery4 findings
- Can Knowledge Editing Really Correct Hallucinations?4 findings
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory4 findings
- Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG Alignment4 findings
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark4 findings
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)4 findings
- DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models4 findings
- DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life4 findings
- Discovering Influential Neuron Path in Vision Transformers4 findings
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs4 findings
- Does Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?4 findings
- Episodic Memories Generation and Evaluation Benchmark for Large Language Models4 findings
- FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs4 findings
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"4 findings
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!4 findings
- First-Person Fairness in Chatbots4 findings
- Geometry of Geophysical Knowledge in TerraMind's Latent Space4 findings
- ICLR: In-Context Learning of Representations4 findings
- Improving Instruction-Following in Language Models through Activation Steering4 findings
- Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct4 findings
- Interpreting CLIP's Image Representation via Text-Based Decomposition4 findings
- KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval4 findings
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations4 findings
- LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models4 findings
- Learning Performance-Improving Code Edits4 findings
- Linear Representations of Political Perspective Emerge in Large Language Models4 findings
- MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models4 findings
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos4 findings
- MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models4 findings
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use4 findings
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models4 findings
- Monitoring Latent World States in Language Models with Propositional Probes4 findings
- More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness4 findings
- Not All Language Model Features Are One-Dimensionally Linear4 findings
- Number Cookbook: Number Understanding of Language Models and How to Improve It4 findings
- On Calibration of LLM-based Guard Models for Reliable Content Moderation4 findings
- On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection4 findings
- On the Role of Attention Heads in Large Language Model Safety4 findings
- One Hundred Neural Networks and Brains Watching Videos: Lessons from Alignment4 findings
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models4 findings
- ReCogLab: a framework testing relational reasoning & cognitive hypotheses on LLMs4 findings
- Revisiting In-context Learning Inference Circuit in Large Language Models4 findings
- Robotouille: An Asynchronous Planning Benchmark for LLM Agents4 findings
- SKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models4 findings
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal4 findings
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents4 findings
- ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents4 findings
- Structure Language Models for Protein Conformation Generation4 findings
- Successor Heads: Recurring, Interpretable Attention Heads In The Wild4 findings
- SysBench: Can LLMs Follow System Message?4 findings
- TRACE: Temporal Grounding Video LLM via Causal Event Modeling4 findings
- Talk like a Graph: Encoding Graphs for Large Language Models4 findings
- The "Law'' of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired Modalities4 findings
- The Generative AI Paradox: “What It Can Create, It May Not Understand”4 findings
- The Hidden Language of Diffusion Models4 findings
- To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts4 findings
- Towards Understanding Factual Knowledge of Large Language Models4 findings
- Training-Free Activation Sparsity in Large Language Models4 findings
- Transformer Block Coupling and its Correlation with Generalization in LLMs4 findings
- Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models4 findings
- Understanding In-Context Learning from Repetitions4 findings
- Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and Intervention4 findings
- Unearthing Skill-level Insights for Understanding Trade-offs of Foundation Models4 findings
- VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning4 findings
- Vision-by-Language for Training-Free Compositional Image Retrieval4 findings
- What does the Knowledge Neuron Thesis Have to do with Knowledge?4 findings
- A Benchmark for Learning to Translate a New Language from One Grammar Book3 findings
- AI Sandbagging: Language Models can Strategically Underperform on Evaluations3 findings
- An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels3 findings
- Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions3 findings
- Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics3 findings
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models3 findings
- Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions3 findings
- BadEdit: Backdooring Large Language Models by Model Editing3 findings
- BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks3 findings
- Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World3 findings
- Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image Classification3 findings
- COCO-Periph: Bridging the Gap Between Human and Machine Perception in the Periphery3 findings
- COLLIE: Systematic Construction of Constrained Text Generation Tasks3 findings
- Can LLM-Generated Misinformation Be Detected?3 findings
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs3 findings
- ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities3 findings
- Circuit Component Reuse Across Tasks in Transformer Language Models3 findings
- Compressing LLMs: The Truth is Rarely Pure and Never Simple3 findings
- Confidence Elicitation: A New Attack Vector for Large Language Models3 findings
- Conversational Drug Editing Using Retrieval and Domain Feedback3 findings
- DENEVIL: TOWARDS DECIPHERING AND NAVIGATING THE ETHICAL VALUES OF LARGE LANGUAGE MODELS VIA INSTRUCTION LEARNING3 findings
- Differentially Private Steering for Large Language Model Alignment3 findings
- Discovering Failure Modes of Text-guided Diffusion Models via Adversarial Search3 findings
- Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers3 findings
- Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models3 findings
- Do LLMs have Consistent Values?3 findings
- Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?3 findings
- Do as We Do, Not as You Think: the Conformity of Large Language Models3 findings
- Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting3 findings
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models3 findings
- EFFICIENT JAILBREAK ATTACK SEQUENCES ON LARGE LANGUAGE MODELS VIA MULTI-ARMED BANDIT-BASED CONTEXT SWITCHING3 findings
- EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition3 findings
- Endless Jailbreaks with Bijection Learning3 findings
- FIXLIP: Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions3 findings
- FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs3 findings
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking3 findings
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems3 findings
- Frozen Transformers in Language Models Are Effective Visual Encoder Layers3 findings
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher3 findings
- Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection3 findings
- Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning3 findings
- Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs3 findings
- How do Language Models Bind Entities in Context?3 findings
- INViTE: INterpret and Control Vision-Language Models with Text Explanations3 findings
- In-Context Learning Dynamics with Random Binary Sequences3 findings
- In-Context Learning Learns Label Relationships but Is Not Conventional Learning3 findings
- InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation3 findings
- Interpretable Diffusion via Information Decomposition3 findings
- Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations3 findings
- Interpreting the Second-Order Effects of Neurons in CLIP3 findings
- Is Self-Repair a Silver Bullet for Code Generation?3 findings
- Is Your Video Language Model a Reliable Judge?3 findings
- JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention3 findings
- Knowledge Localization: Mission Not Accomplished? Enter Query Localization!3 findings
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset3 findings
- Language Modeling Is Compression3 findings
- Language Models Are Implicitly Continuous3 findings
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning3 findings
- Large Language Models Cannot Self-Correct Reasoning Yet3 findings
- Lawma: The Power of Specialization for Legal Annotation3 findings
- Linearity of Relation Decoding in Transformer Language Models3 findings
- LitCab: Lightweight Language Model Calibration over Short- and Long-form Responses3 findings
- Look Before You Leap: Universal Emergent Mechanism for Retrieval in Language Models3 findings
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback3 findings
- Measuring Vision-Language STEM Skills of Neural Models3 findings
- Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models3 findings
- NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language Models3 findings
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?3 findings
- OctoPack: Instruction Tuning Code Large Language Models3 findings
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs3 findings
- On the self-verification limitations of large language models on reasoning and planning tasks3 findings
- OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?3 findings
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization3 findings
- PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts3 findings
- Perturbation-Restrained Sequential Model Editing3 findings
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement3 findings
- Physics of Language Models: Part 3.2, Knowledge Manipulation3 findings
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models3 findings
- Protein Language Model Fitness is a Matter of Preference3 findings
- Quantifying Generalization Complexity for Large Language Models3 findings
- Refining CLIP's Spatial Awareness: A Visual-Centric Perspective3 findings
- Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models3 findings
- Robust LLM safeguarding via refusal feature adversarial training3 findings
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis3 findings
- SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution3 findings
- Safety Alignment Should be Made More Than Just a Few Tokens Deep3 findings
- ScImage: How good are multimodal large language models at scientific text-to-image generation?3 findings
- Scaling and evaluating sparse autoencoders3 findings
- Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification3 findings
- Sparse Autoencoders Find Highly Interpretable Features in Language Models3 findings
- Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models3 findings
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models3 findings
- Supervised Knowledge Makes Large Language Models Better In-context Learners3 findings
- Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs3 findings
- The 3D-PC: a benchmark for visual perspective taking in humans and machines3 findings
- The Alignment Problem from a Deep Learning Perspective3 findings
- The Cost of Scaling Down Large Language Models: Reducing Model Size Affects Memory before In-context Learning3 findings
- The Devil is in the Object Boundary: Towards Annotation-free Instance Segmentation using Foundation Models3 findings
- The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling3 findings
- Towards Interpreting Visual Information Processing in Vision-Language Models3 findings
- Training on the Test Task Confounds Evaluation and Emergence3 findings
- Uncovering Latent Memories in Large Language Models3 findings
- Understanding Catastrophic Forgetting in Language Models via Implicit Inference3 findings
- Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions3 findings
- Understanding and Enhancing the Transferability of Jailbreaking Attacks3 findings
- UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition3 findings
- Vector-ICL: In-context Learning with Continuous Vector Representations3 findings
- VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models3 findings
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning3 findings
- What Makes Large Language Models Reason in (Multi-Turn) Code Generation?3 findings
- What does the TerraMind model look at and think about?3 findings
- Where and Why TerraMind's Optical-to-SAR Generation Fails3 findings
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild3 findings
- WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models3 findings
- A Geometric Framework for Understanding Memorization in Generative Models2 findings
- AI as Humanity’s Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text2 findings
- Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps2 findings
- AutoBencher: Towards Declarative Benchmark Construction2 findings
- Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution2 findings
- Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability2 findings
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs2 findings
- Bisimulation Metric for Model Predictive Control2 findings
- BooookScore: A systematic exploration of book-length summarization in the era of LLMs2 findings
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing2 findings
- Calibrating Expressions of Certainty2 findings
- Can In-context Learning Really Generalize to Out-of-distribution Tasks?2 findings
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks2 findings
- Can We Talk Models Into Seeing the World Differently?2 findings
- Capability Localization: Capabilities Can be Localized rather than Individual Knowledge2 findings
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation2 findings
- Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference2 findings
- Closing the Curious Case of Neural Text Degeneration2 findings
- Competition Dynamics Shape Algorithmic Phases of In-Context Learning2 findings
- Concept vectors in TerraMind latent space2 findings
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning2 findings
- Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model2 findings
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion2 findings
- DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image2 findings
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation2 findings
- Debiasing Algorithm through Model Adaptation2 findings
- Dissecting learning and forgetting in language model finetuning2 findings
- Do LLMs ``know'' internally when they follow instructions?2 findings
- Does Refusal Training in LLMs Generalize to the Past Tense?2 findings
- Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?2 findings
- Efficient Streaming Language Models with Attention Sinks2 findings
- Eliminating Position Bias of Language Models: A Mechanistic Approach2 findings
- Evaluating Large Language Models through Role-Guide and Self-Reflection: A Comparative Study2 findings
- Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?2 findings
- Faithful and Efficient Explanations for Neural Networks via Neural Tangent Kernel Surrogate Models2 findings
- Function Vectors in Large Language Models2 findings
- GAIA: a benchmark for General AI Assistants2 findings
- GNNX-BENCH: Unravelling the Utility of Perturbation-based GNN Explainers through In-depth Benchmarking2 findings
- How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?2 findings
- Improving Reasoning Performance in Large Language Models via Representation Engineering2 findings
- Injecting Universal Jailbreak Backdoors into LLMs in Minutes2 findings
- Input Space Mode Connectivity in Deep Neural Networks2 findings
- InterpGNN: Understand and Improve Generalization Ability of Transdutive GNNs through the Lens of Interplay between Train and Test Nodes2 findings
- Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology2 findings
- Is Factuality Enhancement a Free Lunch For LLMs? Better Factuality Can Lead to Worse Context-Faithfulness2 findings
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching2 findings
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models2 findings
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks2 findings
- Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment2 findings
- Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition2 findings
- LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension2 findings
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models2 findings
- Language Model Cascades: Token-Level Uncertainty And Beyond2 findings
- Language Models Represent Space and Time2 findings
- Large Language Models Assume People are More Rational than We Really are2 findings
- Large Language Models as Automated Aligners for benchmarking Vision-Language Models2 findings
- Large Language Models can Become Strong Self-Detoxifiers2 findings
- Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding2 findings
- Layer-wise linear mode connectivity2 findings
- Learning Dynamics of LLM Finetuning2 findings
- Lines of Thought in Large Language Models2 findings
- Localizing and Editing Knowledge In Text-to-Image Generative Models2 findings
- MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models2 findings
- Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching2 findings
- Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse2 findings
- Memorization Capacity of Multi-Head Attention in Transformers2 findings
- Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity2 findings
- Monet: Mixture of Monosemantic Experts for Transformers2 findings
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning2 findings
- NL-Eye: Abductive NLI For Images2 findings
- Network Memory Footprint Compression Through Jointly Learnable Codebooks and Mappings2 findings
- Neuron based Personality Trait Induction in Large Language Models2 findings
- On Linear Representations and Pretraining Data Frequency in Language Models2 findings
- Parsing neural dynamics with infinite recurrent switching linear dynamical systems2 findings
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting2 findings
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models2 findings
- Residual Stream Analysis with Multi-Layer SAEs2 findings
- Rethinking the Benefits of Steerable Features in 3D Equivariant Graph Neural Networks2 findings
- Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks2 findings
- RocketEval: Efficient automated LLM evaluation via grading checklist2 findings
- SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models2 findings
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions2 findings
- Sparse autoencoders reveal selective remapping of visual concepts during adaptation2 findings
- Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking2 findings
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs2 findings
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems2 findings
- Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling2 findings
- Teaching Large Language Models to Self-Debug2 findings
- The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities2 findings
- The Trickle-down Impact of Reward Inconsistency on RLHF2 findings
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning2 findings
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving2 findings
- Towards Neural Scaling Laws for Time Series Foundation Models2 findings
- Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control2 findings
- Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures2 findings
- Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity2 findings
- Turning large language models into cognitive models2 findings
- Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron2 findings
- Understanding prompt engineering may not require rethinking generalization2 findings
- Unveiling the Pitfalls of Knowledge Editing for Large Language Models2 findings
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks2 findings
- Vanishing Gradients in Reinforcement Finetuning of Language Models2 findings
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models2 findings
- Vision Transformers Need Registers2 findings
- Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations2 findings
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning2 findings
- What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models2 findings
- When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations2 findings
- Where Are We Now? Investigating Spatial Information in TerraMind's Latent Space2 findings
- Your Weak LLM is Secretly a Strong Teacher for Alignment2 findings
- Zero and Few-shot Semantic Parsing with Ambiguous Inputs2 findings
- h4rm3l: A Language for Composable Jailbreak Attack Synthesis2 findings
- $\mathcal{B}$-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis1 finding
- A Formal Framework for Understanding Length Generalization in Transformers1 finding
- A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis1 finding
- A Theoretical Analysis of Self-Supervised Learning for Vision Transformers1 finding
- A Theory for Token-Level Harmonization in Retrieval-Augmented Generation1 finding
- A path-norm toolkit for modern networks: consequences, promises and challenges1 finding
- ADIFF: Explaining audio difference using natural language1 finding
- AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation1 finding
- API Pack: A Massive Multi-Programming Language Dataset for API Call Generation1 finding
- AlpaGasus: Training a Better Alpaca with Fewer Data1 finding
- Alt-Text with Context: Improving Accessibility for Images on Twitter1 finding
- Are Bert Family Good Instruction Followers? A Study on Their Potential And Limitations1 finding
- Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models1 finding
- At Which Training Stage Does Code Data Help LLMs Reasoning?1 finding
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models1 finding
- Bayesian Regularization of Latent Representation1 finding
- Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge1 finding
- Beyond single neurons: population response geometry in digital twins of mouse visual cortex1 finding
- Black-Box Detection of Language Model Watermarks1 finding
- BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks1 finding
- Brain decoding: toward real-time reconstruction of visual perception1 finding
- Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization1 finding
- Circumventing Concept Erasure Methods For Text-To-Image Generative Models1 finding
- Concept Bottleneck Generative Models1 finding
- Concept Bottleneck Language Models For Protein Design1 finding
- Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR Data1 finding
- Context-Parametric Inversion: Why Instruction Finetuning May Not Actually Improve Context Reliance1 finding
- Contrastive Learning is Spectral Clustering on Similarity Graph1 finding
- Controllable Context Sensitivity and the Knob Behind It1 finding
- Cross-Entropy Is All You Need To Invert the Data Generating Process1 finding
- Data-independent Module-aware Pruning for Hierarchical Vision Transformers1 finding
- Decoding Natural Images from EEG for Object Recognition1 finding
- Defining and extracting generalizable interaction primitives from DNNs1 finding
- Demystifying Embedding Spaces using Large Language Models1 finding
- Density estimation with LLMs: a geometric investigation of in-context learning trajectories1 finding
- Detecting Backdoor Samples in Contrastive Language Image Pretraining1 finding
- Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient1 finding
- Does CLIP’s generalization performance mainly stem from high train-test similarity?1 finding
- Does Writing with Language Models Reduce Content Diversity?1 finding
- Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification1 finding
- Efficient Dynamics Modeling in Interactive Environments with Koopman Theory1 finding
- Elucidating the Exposure Bias in Diffusion Models1 finding
- Elucidating the design space of classifier-guided diffusion generation1 finding
- Emergence of a High-Dimensional Abstraction Phase in Language Transformers1 finding
- Enhancing Transferable Adversarial Attacks on Vision Transformers through Gradient Normalization Scaling and High-Frequency Adaptation1 finding
- Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors1 finding
- Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation1 finding
- Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval1 finding
- Failures to Find Transferable Image Jailbreaks Between Vision-Language Models1 finding
- Faithful Rule Extraction for Differentiable Rule Learning Models1 finding
- FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models1 finding
- Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks1 finding
- FreSh: Frequency Shifting for Accelerated Neural Representation Learning1 finding
- From Bricks to Bridges: Product of Invariances to Enhance Latent Space Communication1 finding
- GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt Optimizers1 finding
- Generative Verifiers: Reward Modeling as Next-Token Prediction1 finding
- Gramian Multimodal Representation Learning and Alignment1 finding
- HaDeMiF: Hallucination Detection and Mitigation in Large Language Models1 finding
- Harnessing Webpage UIs for Text-Rich Visual Understanding1 finding
- How new data permeates LLM knowledge and how to dilute it1 finding
- How to Fine-Tune Vision Models with SGD1 finding
- Human-Aligned Chess With a Bit of Search1 finding
- HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks1 finding
- ImpScore: A Learnable Metric For Quantifying The Implicitness Level of Sentences1 finding
- Improving Complex Reasoning with Dynamic Prompt Corruption: A Soft Prompt Optimization Approach1 finding
- Improving Long-Text Alignment for Text-to-Image Diffusion Models1 finding
- Improving protein optimization with smoothed fitness landscapes1 finding
- In Search of Forgotten Domain Generalization1 finding
- Is Large-scale Pretraining the Secret to Good Domain Generalization?1 finding
- Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability1 finding
- LLM-based Typed Hyperresolution for Commonsense Reasoning with Knowledge Bases1 finding
- LLM-grounded Video Diffusion Models1 finding
- Label-Focused Inductive Bias over Latent Object Features in Visual Classification1 finding
- Language-Informed Visual Concept Learning1 finding
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment1 finding
- LeanAgent: Lifelong Learning for Formal Theorem Proving1 finding
- Learning LLM-as-a-Judge for Preference Alignment1 finding
- Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents1 finding
- Locality Alignment Improves Vision-Language Models1 finding
- Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference1 finding
- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations1 finding
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge1 finding
- Massive Editing for Large Language Models via Meta Learning1 finding
- Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers1 finding
- Mechanistic Permutability: Match Features Across Layers1 finding
- Multi-granularity Correspondence Learning from Long-term Noisy Videos1 finding
- NetInfoF Framework: Measuring and Exploiting Network Usable Information1 finding
- NeurFlow: Interpreting Neural Networks through Neuron Groups and Functional Interactions1 finding
- Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization1 finding
- On Evaluating the Durability of Safeguards for Open-Weight LLMs1 finding
- On the Foundations of Shortcut Learning1 finding
- On the Hölder Stability of Multiset and Graph Neural Networks1 finding
- On the Variance of Neural Network Training with respect to Test Sets and Distributions1 finding
- Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation1 finding
- Overthinking the Truth: Understanding how Language Models Process False Demonstrations1 finding
- PRIME: Prioritizing Interpretability in Failure Mode Extraction1 finding
- Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models1 finding
- PlaSma: Procedural Knowledge Models for Language-based Planning and Re-Planning1 finding
- Poison-splat: Computation Cost Attack on 3D Gaussian Splatting1 finding
- Post-hoc Reward Calibration: A Case Study on Length Bias1 finding
- Precise Parameter Localization for Textual Generation in Diffusion Models1 finding
- Predicting Emergent Abilities with Infinite Resolution Evaluation1 finding
- Predictive, scalable and interpretable knowledge tracing on structured domains1 finding
- Probabilistic Adaptation of Black-Box Text-to-Video Models1 finding
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models1 finding
- Provably Safeguarding a Classifier from OOD and Adversarial Samples1 finding
- QPM: Discrete Optimization for Globally Interpretable Image Classification1 finding
- Quantifying the Plausibility of Context Reliance in Neural Machine Translation1 finding
- RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering1 finding
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches1 finding
- Ranking-aware adapter for text-driven image ordering with CLIP1 finding
- ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability1 finding
- Reassessing How to Compare and Improve the Calibration of Machine Learning Models1 finding
- Reclaiming the Source of Programmatic Policies: Programmatic versus Latent Spaces1 finding
- Reconstructive Visual Instruction Tuning1 finding
- Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition1 finding
- Rethinking Self-Distillation: Label Averaging and Enhanced Soft Label Refinement with Partial Labels1 finding
- Retrieval Head Mechanistically Explains Long-Context Factuality1 finding
- Retrieval meets Long Context Large Language Models1 finding
- Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?1 finding
- SWE-bench: Can Language Models Resolve Real-world Github Issues?1 finding
- SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation1 finding
- Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement1 finding
- Simple is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation1 finding
- Skip-Attention: Improving Vision Transformers by Paying Less Attention1 finding
- Stable Segment Anything Model1 finding
- Tailoring Self-Rationalizers with Multi-Reward Distillation1 finding
- Teaching Arithmetic to Small Transformers1 finding
- Temporal Reasoning Transfer from Text to Video1 finding
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models1 finding
- The False Promise of Imitating Proprietary Language Models1 finding
- The Geometry of Categorical and Hierarchical Concepts in Large Language Models1 finding
- The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions1 finding
- The Reasonableness Behind Unreasonable Translation Capability of Large Language Model1 finding
- The Reversal Curse: LLMs trained on “A is B” fail to learn “B is A”1 finding
- TiC-CLIP: Continual Training of CLIP Models1 finding
- To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data Contamination1 finding
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs1 finding
- Toward Understanding In-context vs. In-weight Learning1 finding
- Towards 3D Molecule-Text Interpretation in Language Models1 finding
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods1 finding
- Towards Certification of Uncertainty Calibration under Adversarial Attacks1 finding
- Towards Robust Multi-Modal Reasoning via Model Selection1 finding
- Towards Understanding Sycophancy in Language Models1 finding
- Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias1 finding
- Transformers Learn Low Sensitivity Functions: Investigations and Implications1 finding
- TypedThinker: Diversify Large Language Model Reasoning with Typed Thinking1 finding
- U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models1 finding
- Uncovering Gaps in How Humans and LLMs Interpret Subjective Language1 finding
- Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression1 finding
- Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks1 finding
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial Attacks1 finding
- Unveiling and Manipulating Prompt Influence in Large Language Models1 finding
- Unveiling the Unseen: Identifiable Clusters in Trained Depthwise Convolutional Kernels1 finding
- What Makes a Good Prune? Maximal Unstructured Pruning for Maximal Cosine Similarity1 finding
- Where We Have Arrived in Proving the Emergence of Sparse Interaction Primitives in DNNs1 finding
- WildChat: 1M ChatGPT Interaction Logs in the Wild1 finding
- YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-Dictionary1 finding
- YaRN: Efficient Context Window Extension of Large Language Models1 finding