COLM 2026

Accepted Papers

Filter at the top, pick a paper on the left, read on the right. Room colours match the schedule. Copy the URL to share a filtered view or a single paper.

856 of 856 papers
Day
Session
Room
Type
Topic
  • Asymmetric Idiosyncrasies in Multimodal Models

    Tue Oct 6Imperial Ballroom #1Multimodal LMs
  • Source-Modality Monitoring in Vision-Language Models

    ORALTue Oct 6Imperial Ballroom #2Multimodal LMs
  • When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

    Tue Oct 6Imperial Ballroom #3All about safety
  • V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding--Refusal Coupling Failure

    Tue Oct 6Imperial Ballroom #4All about safety
  • Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions

    Tue Oct 6Imperial Ballroom #5All about safety
  • Most current model organisms are leaky: Perplexity differencing often reveals finetuning objectives

    Tue Oct 6Imperial Ballroom #6All about safety
  • A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

    Tue Oct 6Imperial Ballroom #7All about safety
  • Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

    Tue Oct 6Imperial Ballroom #8Inference algorithms for LMs
  • Fragility Under Pressure: Evaluating the Iterative Stability of LLMs in Constrained Interactive Coding

    Tue Oct 6Imperial Ballroom #9All about evaluation
  • LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling

    Tue Oct 6Imperial Ballroom #10All about evaluation
  • Are LLMs Robust Enough for Operations Research? An Adversarial Attack and Defense Study via AttackOR

    Tue Oct 6Imperial Ballroom #11All about safety
  • ConformalGuard: False-Positive-Controlled Action Gating for LM Classifiers

    Tue Oct 6Imperial Ballroom #12All about evaluation
  • BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

    Tue Oct 6Imperial Ballroom #13All about evaluation
  • Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

    Tue Oct 6Imperial Ballroom #14All about evaluation
  • FERA: Uncertainty-Aware Federated Reasoning for Large Language Models

    Tue Oct 6Imperial Ballroom #15Learning algorithms for LMs
  • KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs

    Tue Oct 6Imperial Ballroom #16LMs for everyone
  • Learning to Learn from Language Feedback with Social Meta-Learning

    Tue Oct 6Imperial Ballroom #17LMs and interactions
  • In-Context Exploration-Exploitation is Biased by Semantic Priors

    Tue Oct 6Imperial Ballroom #18Mind, brain, philosophy, law & LMs
  • ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents

    Tue Oct 6Imperial Ballroom #19LMs with tools and code
  • Latent Context Compilation: Distilling Long Context into Compact Portable Memory

    Tue Oct 6Imperial Ballroom #20Compute-efficient LMs
  • A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

    Tue Oct 6Imperial Ballroom #21Compute-efficient LMs
  • HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

    Tue Oct 6Imperial Ballroom #22Compute-efficient LMs
  • Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

    Tue Oct 6Imperial Ballroom #23Compute-efficient LMs
  • Learning to Refer from Estimated Listener Gaze

    Tue Oct 6Imperial Ballroom #24Multimodal LMs
  • Improving GUI Grounding with Explicit Position-to-Coordinate Mapping

    Tue Oct 6Imperial Ballroom #25LMs with tools and code
  • v1: Learning to Point Visual Tokens for Multimodal Mathematical Grounded Reasoning

    Tue Oct 6Imperial Ballroom #26Multimodal LMs
  • To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning

    Tue Oct 6Imperial Ballroom #27Science of LMs
  • Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs

    Tue Oct 6Imperial Ballroom #28LMs and the world
  • Reasoning over Non-Parametric Memory in Multilingual LLMs: A Cognitive-Based Analysis of Relational Knowledge

    Tue Oct 6Imperial Ballroom #29LMs for everyone
  • Uncovering the computational ingredients that support human-like conceptual representations in large language models

    Tue Oct 6Imperial Ballroom #30Mind, brain, philosophy, law & LMs
  • Large Language Models Align with the Human Brain during Creative Thinking

    Tue Oct 6Imperial Ballroom #31Mind, brain, philosophy, law & LMs
  • CreativityBench: Evaluating Creative Reasoning via Affordance-Based Tool Repurposing

    Tue Oct 6Imperial Ballroom #32All about evaluation
  • The Tool Illusion: Rethinking Tool Use in Web Agents

    Tue Oct 6Imperial Ballroom #33LMs with tools and code
  • CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

    ORALTue Oct 6Imperial Ballroom #34LMs and interactions
  • SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

    Tue Oct 6Imperial Ballroom #35Learning algorithms for LMs
  • Characterizing Model-Native Skills

    Tue Oct 6Imperial Ballroom #36Science of LMs
  • Capability Provenance in Language Models: A Case Study in Social Reasoning

    Tue Oct 6Imperial Ballroom #37All about data
  • Prototype Language Models

    Tue Oct 6Imperial Ballroom #38Science of LMs
  • Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

    Tue Oct 6Imperial Ballroom #39All about training
  • Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

    Tue Oct 6Imperial Ballroom #40All about training
  • Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

    Tue Oct 6Imperial Ballroom #41Learning algorithms for LMs
  • MoRE: Mixture of Reused Experts

    Tue Oct 6Imperial Ballroom #42Learning algorithms for LMs
  • Do Depth-Grown Models Overcome The Curse Of Depth? An In-Depth Analysis

    Tue Oct 6Imperial Ballroom #43Learning algorithms for LMs
  • Parcae: Scaling Laws For Stable Looped Language Models

    Tue Oct 6Imperial Ballroom #44Science of LMs
  • A Spectral Transport Mechanism Underlying the Training Dynamics of Large Language Models

    Tue Oct 6Imperial Ballroom #45Science of LMs
  • MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU

    Tue Oct 6Imperial Ballroom #46Engineering for large LMs
  • KronQ: LLM Quantization via Kronecker-Factored Hessian

    Tue Oct 6Imperial Ballroom #47Compute-efficient LMs
  • Covariance-Aware Transformers for Quadratic Programming and Decision Making

    Tue Oct 6Imperial Ballroom #48Diverse domains & novel applications
  • Squeeze Evolve: A Unified Multi-Model Orchestration Framework for Verifier-Free Evolution

    Tue Oct 6Imperial Ballroom #49Engineering for large LMs
  • DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent

    Tue Oct 6Imperial Ballroom #50Inference algorithms for LMs
  • Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification

    Tue Oct 6Imperial Ballroom #51Inference algorithms for LMs
  • Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving

    Tue Oct 6Imperial Ballroom #52All about training
  • Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models

    Tue Oct 6Imperial Ballroom #53All about training
  • CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

    Tue Oct 6Imperial Ballroom #54All about training
  • Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

    Tue Oct 6Imperial Ballroom #55LMs with tools and code
  • Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification

    Tue Oct 6Imperial Ballroom #56Multimodal LMs
  • Reasoning about Intent for Ambiguous Requests

    ORALTue Oct 6Imperial Ballroom #57LMs and interactions
  • Knowing Before Answering: Decoding Language Models for Reliable RAG

    Tue Oct 6Imperial Ballroom #58LMs and the world
  • Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines

    Tue Oct 6Imperial Ballroom #59Inference algorithms for LMs
  • RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

    Tue Oct 6Imperial Ballroom #60Multimodal LMs
  • Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

    Tue Oct 6Imperial Ballroom #61All about evaluation
  • Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction

    Tue Oct 6Imperial Ballroom #62LMs and interactions
  • PreScam: A Benchmark for Predicting Scam Progression from Early Conversations

    Tue Oct 6Imperial Ballroom #63All about safety
  • Mind the Sim2Real Gap in User Simulation for Agentic Tasks

    Tue Oct 6Imperial Ballroom #64All about evaluation
  • InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation

    Tue Oct 6Imperial Ballroom #65LMs and the world
  • Identifying Introspection From the Inside

    Tue Oct 6Imperial Ballroom #66Mind, brain, philosophy, law & LMs
  • Rule vs. Consequence: Dissociable Internal Representations of Moral Reasoning in LLMs

    Tue Oct 6Imperial Ballroom #67Mind, brain, philosophy, law & LMs
  • Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

    Tue Oct 6Imperial Ballroom #68LMs and interactions
  • Misalignment Contagion: Can a Misaligned Minority Shift Aligned Agents in Multi-Agent LLM Deliberation?

    Tue Oct 6Imperial Ballroom #69All about safety
  • Deep Research Agents Bring Deeper Harm

    Tue Oct 6Imperial Ballroom #70All about safety
  • Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs

    Tue Oct 6Imperial Ballroom #71All about safety
  • The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

    Tue Oct 6Imperial Ballroom #72LMs and interactions
  • Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training

    Tue Oct 6Imperial Ballroom #73LMs and interactions
  • Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

    Tue Oct 6Franciscan A #74LMs and interactions
  • Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

    Tue Oct 6Franciscan A #75All about safety
  • Context-based normative simulacra from narrative fiction

    Tue Oct 6Franciscan A #76All about safety
  • Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

    Tue Oct 6Franciscan A #77All about safety
  • Subliminal Steering: Encoding and Detecting Hidden Signals

    Tue Oct 6Franciscan A #78All about safety
  • Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior

    Tue Oct 6Franciscan A #79Science of LMs
  • Activation Steering via Generative Causal Mediation

    Tue Oct 6Franciscan A #80Science of LMs
  • Steering LLMs for Culturally Localized Generation

    Tue Oct 6Franciscan B #81LMs for everyone
  • MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

    Tue Oct 6Franciscan B #82LMs for everyone
  • Evaluating and Mitigating Misgendering in English-to-Hindi Machine Translation

    Tue Oct 6Franciscan B #83LMs for everyone
  • MMMG: A Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

    Tue Oct 6Franciscan B #84Multimodal LMs
  • SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

    Tue Oct 6Franciscan B #85Multimodal LMs
  • Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities

    Tue Oct 6Franciscan B #86All about evaluation
  • Valley3: Scaling Omni Foundation Models for E-commerce

    Tue Oct 6Franciscan B #87Diverse domains & novel applications
  • TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

    Tue Oct 6Franciscan B #88Multimodal LMs
  • FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

    Tue Oct 6Franciscan B #89Compute-efficient LMs
  • FLINT: Influence-Guided Active Learning Framework for LoRA via Curvature-Aware Data Selection and Fine-Tuning

    Tue Oct 6Franciscan B #90All about data
  • Towards Active Synthetic Data Generation for Finetuning Language Models

    Tue Oct 6Franciscan C #91All about data
  • Synthetic Data for any Differentiable Target

    Tue Oct 6Franciscan C #92All about data
  • Synthesizing Instruction-Tuning Datasets with Contrastive Decoding

    Tue Oct 6Franciscan C #93All about data
  • From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

    Tue Oct 6Franciscan C #94Multimodal LMs
  • All-Weather VLM: Enhancing Vision Language Models' Robustness Under Adverse Imaging Conditions

    Tue Oct 6Franciscan C #95Multimodal LMs
  • Wiener Filtering for VLM Hallucination Suppression

    Tue Oct 6Franciscan C #96Multimodal LMs
  • Why Fine-Tuning Encourages Hallucinations and How to Fix It

    Tue Oct 6Franciscan C #97All about training
  • Fine-Tuning Was Not Broken but Diffuse: Precise Knowledge Editing via Projected Error Signals

    Tue Oct 6Franciscan C #98All about training
  • LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance

    Tue Oct 6Franciscan C #99LMs and the world
  • When Verification Fails: How Compositionally Infeasible Claims Escape Rejection

    Tue Oct 6Franciscan C #100LMs and the world
  • ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability

    Tue Oct 6Grand Ballroom #101Inference algorithms for LMs
  • Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces

    Tue Oct 6Grand Ballroom #102All about evaluation
  • Understanding and Mitigating Premature Confidence for Better LLM Reasoning

    Tue Oct 6Grand Ballroom #103Inference algorithms for LMs
  • Counterfactual Simulation Training for Chain-of-Thought Faithfulness

    Tue Oct 6Grand Ballroom #104Science of LMs
  • Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning

    Tue Oct 6Grand Ballroom #105All about data
  • Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

    Tue Oct 6Grand Ballroom #106All about data
  • PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost

    Tue Oct 6Grand Ballroom #107All about training
  • V_{0.5}: Generalist Value Model as a Prior for Sparse RL Rollouts

    Tue Oct 6Grand Ballroom #108All about training
  • HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning

    Tue Oct 6Grand Ballroom #109All about training
  • Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking

    Tue Oct 6Grand Ballroom #110All about training
  • Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning

    Tue Oct 6Grand Ballroom #111All about safety
  • A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges

    Tue Oct 6Grand Ballroom #112All about evaluation
  • FLY-EVAL++: Evidence-Based Evaluation for Safety-Constrained Modeling in Embodied Systems

    Tue Oct 6Grand Ballroom #113LMs and embodiment
  • Measuring Nines of Reliability: Sample-Efficient LLM Evaluation in Saturated GSM Benchmarks

    Tue Oct 6Grand Ballroom #114All about evaluation
  • The Illusion of Stochasticity in LLMs

    Tue Oct 6Grand Ballroom #115Inference algorithms for LMs
  • Why Gaussian Diffusion Models Fail on Discrete Data?

    Tue Oct 6Grand Ballroom #116Learning algorithms for LMs
  • Measuring AI "Slop" in Text

    Tue Oct 6Grand Ballroom #117All about evaluation
  • Making Grid Beam Search Less Greedy

    Tue Oct 6Grand Ballroom #118Inference algorithms for LMs
  • Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

    Tue Oct 6Grand Ballroom #119Compute-efficient LMs
  • More Than Words: Compositional Tokenization for Efficient Language Models

    ORALTue Oct 6Grand Ballroom #120Compute-efficient LMs
  • Inference Scaling of LLM Ensembling: Bridging Token Spaces with Token Translation

    Tue Oct 6Grand Ballroom #121Inference algorithms for LMs
  • Tiny Aya: Bridging Scale and Multilingual Depth

    Tue Oct 6Grand Ballroom #122LMs for everyone
  • MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    Tue Oct 6Grand Ballroom #123LMs for everyone
  • The Embedder's Dilemma: LLMs Are Better, but at What Cost?

    Tue Oct 6Grand Ballroom #124Compute-efficient LMs
  • LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

    Tue Oct 6Grand Ballroom #125Inference algorithms for LMs
  • Scaling Test-Time Compute for Agentic Coding

    Tue Oct 6Grand Ballroom #126Inference algorithms for LMs
  • SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

    Tue Oct 6Grand Ballroom #127LMs with tools and code
  • Learning Reasoning World Models for Parallel Code

    Tue Oct 6Grand Ballroom #128LMs with tools and code
  • SuperCoder: Assembly Program Superoptimization with Large Language Models

    Tue Oct 6Grand Ballroom #129LMs with tools and code
  • DarwinLM: Evolutionary Structured Pruning of Large Language Models

    Tue Oct 6Grand Ballroom #130Compute-efficient LMs
  • Sememe: Causal Modeling for Task-Aware Depth Pruning of Large Language Models

    Tue Oct 6Grand Ballroom #131Compute-efficient LMs
  • SEER: Long-Context Reasoning via Selective Visual-Text Compression

    Tue Oct 6Grand Ballroom #132Compute-efficient LMs
  • Resa: Efficient Reasoning Models via SAEs

    Tue Oct 6Grand Ballroom #133Compute-efficient LMs
  • π^2: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models

    Tue Oct 6Grand Ballroom #134All about data
  • LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

    Tue Oct 6Grand Ballroom #135Multimodal LMs
  • OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning

    Tue Oct 6Grand Ballroom #136LMs with tools and code
  • Co-Evolving Structured Knowledge and Reasoning in Language Models

    Tue Oct 6Grand Ballroom #137LMs and the world
  • SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning

    Tue Oct 6Grand Ballroom #138LMs with tools and code
  • CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

    Tue Oct 6Grand Ballroom #139LMs with tools and code
  • RecaLLM: Addressing the Lost-in-Thought Phenomenon with Explicit In-Context Retrieval

    Tue Oct 6Grand Ballroom #140LMs and the world
  • Unable to Forget: Proactive Interference Reveals Working Memory Limits in LLMs Beyond Context Length

    Tue Oct 6Grand Ballroom #141Mind, brain, philosophy, law & LMs
  • Does Episodic Memory Help Close the Lexical Frequency Gap in Sensitivity to Syntactic Contrasts? A Test Using Retrieval-Augmented Language Models

    Tue Oct 6Grand Ballroom #142Mind, brain, philosophy, law & LMs
  • Combee: Scaling Parallel Prompt Learning for Self-Improving LLM Agents

    Tue Oct 6Imperial Ballroom #1All about training
  • Training Proactive and Personalized LLM Agents

    Tue Oct 6Imperial Ballroom #2LMs and interactions
  • Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans

    Tue Oct 6Imperial Ballroom #3LMs and the world
  • Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis

    Tue Oct 6Imperial Ballroom #4LMs and the world
  • Small Foundation Models of Human Cognition and Behaviour

    Tue Oct 6Imperial Ballroom #5Mind, brain, philosophy, law & LMs
  • Superficial Beliefs in LLM Decision-Making

    Tue Oct 6Imperial Ballroom #6Mind, brain, philosophy, law & LMs
  • When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

    ORALTue Oct 6Imperial Ballroom #8Mind, brain, philosophy, law & LMs
  • When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don’t

    Tue Oct 6Imperial Ballroom #9Mind, brain, philosophy, law & LMs
  • Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs

    Tue Oct 6Imperial Ballroom #10Science of LMs
  • Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

    Tue Oct 6Imperial Ballroom #11Science of LMs
  • ADAG: Automatically Describing Attribution Graphs

    Tue Oct 6Imperial Ballroom #12Science of LMs
  • Data Canvas: A Provenance-Guided Harness for Agentic Data Engineering

    Tue Oct 6Imperial Ballroom #13LMs with tools and code
  • The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

    Tue Oct 6Imperial Ballroom #14LMs with tools and code
  • Recursive Agent Optimization

    Tue Oct 6Imperial Ballroom #15All about training
  • Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning

    Tue Oct 6Imperial Ballroom #16LMs with tools and code
  • SkillFoundry: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

    Tue Oct 6Imperial Ballroom #17LMs with tools and code
  • SkillFlow: Scalable and Efficient Agent Skill Retrieval System

    Tue Oct 6Imperial Ballroom #18LMs with tools and code
  • CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search

    Tue Oct 6Imperial Ballroom #19LMs with tools and code
  • Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    Tue Oct 6Imperial Ballroom #20LMs with tools and code
  • Don’t Lose the Thread: Empowering Long-Horizon LLM Agents with Cognitive Resource Self-Allocation

    Tue Oct 6Imperial Ballroom #21LMs with tools and code
  • AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

    Tue Oct 6Imperial Ballroom #22Inference algorithms for LMs
  • SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

    Tue Oct 6Imperial Ballroom #23Inference algorithms for LMs
  • Failure-Aware Penetration Testing via Typed Failure Tokens: From Calibrated Prediction to Structured Recovery

    Tue Oct 6Imperial Ballroom #24All about safety
  • A Diagnostic Failure Taxonomy and Trajectory Critic for Data Agent Improvement

    Tue Oct 6Imperial Ballroom #25LMs with tools and code
  • MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs

    Tue Oct 6Imperial Ballroom #26Diverse domains & novel applications
  • MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

    Tue Oct 6Imperial Ballroom #27Diverse domains & novel applications
  • JMedEthicBench: A Multi-Turn Adversarial Benchmark for Japanese Medical Ethics Alignment in LLMs

    Tue Oct 6Imperial Ballroom #28Diverse domains & novel applications
  • TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    Tue Oct 6Imperial Ballroom #29All about safety
  • LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X

    Tue Oct 6Imperial Ballroom #30All about data
  • Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality

    Tue Oct 6Imperial Ballroom #31Engineering for large LMs
  • In-context superposition: human-like working memory interference in large language models

    Tue Oct 6Imperial Ballroom #32Mind, brain, philosophy, law & LMs
  • The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

    Tue Oct 6Imperial Ballroom #33Science of LMs
  • Do Sparse Autoencoders Capture Concept Manifolds?

    Tue Oct 6Imperial Ballroom #34Science of LMs
  • Semantic Structure of Feature Space in Large Language Models

    Tue Oct 6Imperial Ballroom #35Science of LMs
  • LLM2Vec-Gen: Generative Embeddings from Large Language Models

    Tue Oct 6Imperial Ballroom #36Learning algorithms for LMs
  • SMETA-ZSL: Semantic Meta-Alignment for Zero-Shot Threat Classification

    Tue Oct 6Imperial Ballroom #37All about safety
  • GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints

    Tue Oct 6Imperial Ballroom #38Learning algorithms for LMs
  • Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning

    Tue Oct 6Imperial Ballroom #39All about safety
  • ContextLeak: Auditing Leakage in Private In-Context Learning Methods

    Tue Oct 6Imperial Ballroom #40All about safety
  • Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes

    Tue Oct 6Imperial Ballroom #41All about evaluation
  • GLiGuard: Schema-Conditioned Classification for LLM Safeguard

    Tue Oct 6Imperial Ballroom #42All about safety
  • Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT

    Tue Oct 6Imperial Ballroom #43All about safety
  • Strong but Brittle: Simple Attacks Subvert Reasoning-based Safety Guardrails

    Tue Oct 6Imperial Ballroom #44All about safety
  • Understanding the Effects of Safety Unalignment on Large Language Models

    Tue Oct 6Imperial Ballroom #45All about safety
  • Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models

    Tue Oct 6Imperial Ballroom #46LMs and the world
  • Diversifying Multiple Generative Agents by Aligning with Human Populations

    Tue Oct 6Imperial Ballroom #47LMs for everyone
  • Can Large Language Models Match Human Diversity in Educational Content Generation?

    Tue Oct 6Imperial Ballroom #48LMs and the world
  • The Garden of Forking Prompts: How Users Explore Narrative Space in Story Generation

    Tue Oct 6Imperial Ballroom #49LMs and interactions
  • Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

    Tue Oct 6Imperial Ballroom #50LMs for everyone
  • Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

    Tue Oct 6Imperial Ballroom #51LMs for everyone
  • Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

    Tue Oct 6Imperial Ballroom #52LMs for everyone
  • Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

    Tue Oct 6Imperial Ballroom #53LMs for everyone
  • English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

    Tue Oct 6Imperial Ballroom #54LMs for everyone
  • MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models

    Tue Oct 6Imperial Ballroom #55LMs for everyone
  • Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

    Tue Oct 6Imperial Ballroom #56Multimodal LMs
  • ComDoc: A Comprehensive Benchmark for Document Retrieval-Augmented Generation

    Tue Oct 6Imperial Ballroom #57LMs and the world
  • Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

    Tue Oct 6Imperial Ballroom #58LMs and the world
  • OSCAR : Orchestrated Self-verification and Cross-path Refinement

    Tue Oct 6Imperial Ballroom #59LMs and the world
  • SALA: Syntax-Aware Logit Adjustment for Open-Ended Text Generation

    Tue Oct 6Imperial Ballroom #60Inference algorithms for LMs
  • Where Does the Relative Clause Attach? Probing Prosodic and Semantic Sensitivity in Large Audio Language Models

    Tue Oct 6Imperial Ballroom #61Mind, brain, philosophy, law & LMs
  • Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages

    Tue Oct 6Imperial Ballroom #62Science of LMs
  • Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs

    Tue Oct 6Imperial Ballroom #63LMs for everyone
  • Verb-ICL: Rethinking In-Context Learning for Structured Prediction

    Tue Oct 6Imperial Ballroom #64All about training
  • RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

    Tue Oct 6Imperial Ballroom #65All about training
  • A Unified Model and Document Representation for On-Device Retrieval-Augmented Generation

    Tue Oct 6Imperial Ballroom #66LMs and the world
  • Back to Basics: Let Conversational Agents Remember with Just Retrieval and Generation

    Tue Oct 6Imperial Ballroom #67LMs and the world
  • What Makes Retrieval Work for Pedagogical Dialogue Annotation? Indexing Granularity over Retriever Adaptation

    Tue Oct 6Imperial Ballroom #68LMs and the world
  • KT4EQG: Personalized Exercise Question Generation via Knowledge Tracing

    Tue Oct 6Imperial Ballroom #69LMs and the world
  • SERUM: State Extraction and Refinement for User Modeling

    Tue Oct 6Imperial Ballroom #70Multimodal LMs
  • EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

    Tue Oct 6Imperial Ballroom #71Multimodal LMs
  • MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

    Tue Oct 6Imperial Ballroom #72Multimodal LMs
  • Explainable AI-Generated Image Detection RewardBench

    Tue Oct 6Imperial Ballroom #73All about safety
  • Reasoning with Image Generation

    Tue Oct 6Franciscan A #74Multimodal LMs
  • FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

    Tue Oct 6Franciscan A #75LMs with tools and code
  • Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation

    Tue Oct 6Franciscan A #76LMs and interactions
  • Effective Strategies for Asynchronous Software Engineering Agents

    Tue Oct 6Franciscan A #77LMs with tools and code
  • HAI-Agent: Improving Long Horizon Software Engineering with Handoff Interventions

    Tue Oct 6Franciscan A #78LMs with tools and code
  • From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents

    Tue Oct 6Franciscan A #79LMs with tools and code
  • MechMath: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving

    Tue Oct 6Franciscan A #80Diverse domains & novel applications
  • LeanGeo: Formalizing Competitional Geometry problems in Lean

    Tue Oct 6Franciscan B #81Diverse domains & novel applications
  • Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

    Tue Oct 6Franciscan B #82Diverse domains & novel applications
  • CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge

    Tue Oct 6Franciscan B #83All about evaluation
  • DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?

    Tue Oct 6Franciscan B #84All about evaluation
  • InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

    Tue Oct 6Franciscan B #85Diverse domains & novel applications
  • MegaScience: Pushing the Frontiers of Open Post-Training Datasets for Science Reasoning

    Tue Oct 6Franciscan B #86All about data
  • LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

    Tue Oct 6Franciscan B #87LMs with tools and code
  • Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks

    Tue Oct 6Franciscan B #88Inference algorithms for LMs
  • FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modular Collaboration

    Tue Oct 6Franciscan B #89Inference algorithms for LMs
  • Test-Time Scaling Makes Overtraining Compute-Optimal

    Tue Oct 6Franciscan B #90Science of LMs
  • Relative Scaling Laws for LLMs

    Tue Oct 6Franciscan C #91Science of LMs
  • Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

    Tue Oct 6Franciscan C #92All about evaluation
  • Cost-Efficient Estimation of General Abilities Across Benchmarks

    Tue Oct 6Franciscan C #93All about evaluation
  • RankBALD: Ranking-Aligned Active Evaluation for Language Models

    Tue Oct 6Franciscan C #94All about evaluation
  • LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights

    Tue Oct 6Franciscan C #95Compute-efficient LMs
  • Hyperloop Transformers

    Tue Oct 6Franciscan C #96Learning algorithms for LMs
  • HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

    Tue Oct 6Franciscan C #97Compute-efficient LMs
  • Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression

    Tue Oct 6Franciscan C #98Compute-efficient LMs
  • QLPO: Quadrant-weighted sampling for Length-aware Policy Optimization

    Tue Oct 6Franciscan C #99Compute-efficient LMs
  • FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization

    Tue Oct 6Franciscan C #100All about training
  • Learn from Zero: Policy Optimization under Vanishing Advantage

    Tue Oct 6Grand Ballroom #101All about training
  • Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR

    Tue Oct 6Grand Ballroom #102All about training
  • From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

    Tue Oct 6Grand Ballroom #103All about training
  • To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models

    Tue Oct 6Grand Ballroom #104All about training
  • Building Multi-Task Agentic LLMs via Two-Phase Distillation

    Tue Oct 6Grand Ballroom #105All about training
  • Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

    Tue Oct 6Grand Ballroom #106All about training
  • Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

    Tue Oct 6Grand Ballroom #107All about training
  • Memorization Dynamics in Knowledge Distillation for Language Models

    Tue Oct 6Grand Ballroom #108Science of LMs
  • Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

    Tue Oct 6Grand Ballroom #109All about safety
  • ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding of Health Information, but at What Cost?

    Tue Oct 6Grand Ballroom #110Diverse domains & novel applications
  • Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs

    Tue Oct 6Grand Ballroom #111LMs and the world
  • LLMs as Content Curators: A Large-Scale Audit of Recommendation Bias Across Providers and Platforms

    Tue Oct 6Grand Ballroom #112All about evaluation
  • Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring

    Tue Oct 6Grand Ballroom #113All about evaluation
  • Investigating Social Bias Changes in Quantized Large Language Models

    Tue Oct 6Grand Ballroom #114Compute-efficient LMs
  • MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization

    Tue Oct 6Grand Ballroom #115Compute-efficient LMs
  • Muon^p: Muon with Fractional Spectral Powers

    Tue Oct 6Grand Ballroom #116Engineering for large LMs
  • Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

    Tue Oct 6Grand Ballroom #117Science of LMs
  • PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy

    Tue Oct 6Grand Ballroom #118Science of LMs
  • Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

    Tue Oct 6Grand Ballroom #119LMs for everyone
  • Overton Pluralistic Reinforcement Learning for Large Language Models

    Tue Oct 6Grand Ballroom #120LMs for everyone
  • Pareto-Optimal RTL Code Generation via Multi-Objective Reinforcement Learning with Large Language Models

    Tue Oct 6Grand Ballroom #121Diverse domains & novel applications
  • Scaling Self-Play with Self-Guidance

    Tue Oct 6Grand Ballroom #122All about training
  • How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

    Tue Oct 6Grand Ballroom #123All about training
  • What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features

    Tue Oct 6Grand Ballroom #124LMs for everyone
  • Capacity-Dependent Effects of Data Selection for Mathematical Reasoning

    Tue Oct 6Grand Ballroom #125All about data
  • (How) Learning Rates Regulate Catastrophic Overtraining

    Tue Oct 6Grand Ballroom #126Science of LMs
  • GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM Post Training

    Tue Oct 6Grand Ballroom #127Science of LMs
  • Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM post-training

    Tue Oct 6Grand Ballroom #128All about training
  • Restoring Generalization in Fine-tuned Multimodal LLMs via Geometric Alignment

    Tue Oct 6Grand Ballroom #129Learning algorithms for LMs
  • The Generalization Ridge: Information Flow in Natural Language Generation

    Tue Oct 6Grand Ballroom #130Science of LMs
  • Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

    Tue Oct 6Grand Ballroom #131Mind, brain, philosophy, law & LMs
  • Barriers to Universal Reasoning with Transformers (and How to Overcome them)

    Tue Oct 6Grand Ballroom #132Science of LMs
  • Algebraic Decomposition Theory for Transformer Length Generalization

    Tue Oct 6Grand Ballroom #133Science of LMs
  • Olmo Hybrid: From Theory to Practice and Back

    Tue Oct 6Grand Ballroom #134Learning algorithms for LMs
  • Cylon: Asynchronous Linear Attention

    Tue Oct 6Grand Ballroom #135Engineering for large LMs
  • Switching Linear Attention

    Tue Oct 6Grand Ballroom #136Learning algorithms for LMs
  • IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

    Tue Oct 6Grand Ballroom #137Compute-efficient LMs
  • Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

    Tue Oct 6Grand Ballroom #138Compute-efficient LMs
  • Faster Superword Tokenization

    Tue Oct 6Grand Ballroom #139Compute-efficient LMs
  • A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Large Language Model Training

    ORALTue Oct 6Grand Ballroom #140Engineering for large LMs
  • Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

    Tue Oct 6Grand Ballroom #141All about data
  • Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement

    Tue Oct 6Grand Ballroom #142Learning algorithms for LMs
  • PromptBridge: Cross-Model Prompt Transfer for Large Language Models

    Tue Oct 6Grand Ballroom #143All about training
  • Phonological Perception of Sign Language Models

    ORALTue Oct 6Grand Ballroom #144Multimodal LMs
  • Attribution Bias in Large Language Models

    ORALTue Oct 6Grand Ballroom #145All about evaluation
  • Decocted Experience Improves Test-Time Inference in LLM Agents

    Wed Oct 7Imperial Ballroom #1Inference algorithms for LMs
  • Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent

    Wed Oct 7Imperial Ballroom #2Diverse domains & novel applications
  • PolicyBank: Evolving Policy Understanding for LLM Agents

    Wed Oct 7Imperial Ballroom #3LMs with tools and code
  • Learning Steerable Clarification Policies with Collaborative Self-play

    Wed Oct 7Imperial Ballroom #4LMs and interactions
  • StoryScope: Investigating idiosyncrasies in AI fiction

    Wed Oct 7Imperial Ballroom #5Diverse domains & novel applications
  • Can LLMs Introspect? A Reality Check

    Wed Oct 7Imperial Ballroom #6Mind, brain, philosophy, law & LMs
  • Do LLMs Benefit From Their Own Words?

    Wed Oct 7Imperial Ballroom #7LMs and interactions
  • How Conversational Pressure Alters Belief in Large Language Models: A Hidden State Boundary Perspective

    Wed Oct 7Imperial Ballroom #8LMs and interactions
  • Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State

    Wed Oct 7Imperial Ballroom #9Diverse domains & novel applications
  • Logarithmic Scores, Power-Law Discoveries: Disentangling Measurement from Coverage in Agent-Based Evaluation

    Wed Oct 7Imperial Ballroom #10All about evaluation
  • The State and Fate of Multilingual, Contextual Evaluation in the NLP World

    Wed Oct 7Imperial Ballroom #11LMs for everyone
  • AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

    Wed Oct 7Imperial Ballroom #12Mind, brain, philosophy, law & LMs
  • TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

    Wed Oct 7Imperial Ballroom #13All about evaluation
  • Benchmarking of Automated Assessment of K-5 Student Narratives Using Large Language Models

    Wed Oct 7Imperial Ballroom #14Diverse domains & novel applications
  • Qworld: Question-Specific Evaluation Criteria for LLMs

    Wed Oct 7Imperial Ballroom #15All about evaluation
  • No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

    Wed Oct 7Imperial Ballroom #16All about evaluation
  • Who Checks the Citations? Benchmarking Legal Hallucination Detection

    Wed Oct 7Imperial Ballroom #17LMs and the world
  • BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

    Wed Oct 7Imperial Ballroom #18LMs and the world
  • Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

    Wed Oct 7Imperial Ballroom #19LMs for everyone
  • TokEval: A Tokenizer Analysis Suite

    Wed Oct 7Imperial Ballroom #20All about evaluation
  • Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

    Wed Oct 7Imperial Ballroom #21Science of LMs
  • Convergent Evolution: How Different Language Models Learn Similar Number Representations

    Wed Oct 7Imperial Ballroom #22Science of LMs
  • Bottom-Up Interpretability of Pretraining Dynamics via Loss Curve Decomposition

    Wed Oct 7Imperial Ballroom #23Science of LMs
  • What Do Language Models Learn and When? The Implicit Curriculum Hypothesis

    ORALWed Oct 7Imperial Ballroom #24Science of LMs
  • Understanding Primacy Effects in Large Language Models with Sparse Autoencoders

    Wed Oct 7Imperial Ballroom #25Science of LMs
  • Differences in Text Generated by Diffusion and Autoregressive Language Models

    Wed Oct 7Imperial Ballroom #26Learning algorithms for LMs
  • Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models

    Wed Oct 7Imperial Ballroom #27Learning algorithms for LMs
  • Cross-Model Disagreement as a Label-Free Correctness Signal

    Wed Oct 7Imperial Ballroom #28All about evaluation
  • Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

    Wed Oct 7Imperial Ballroom #29Inference algorithms for LMs
  • Process Supervision of Confidence Margin for Calibrated LLM Reasoning

    Wed Oct 7Imperial Ballroom #30All about training
  • Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA

    Wed Oct 7Imperial Ballroom #31All about training
  • Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models

    Wed Oct 7Imperial Ballroom #32Multimodal LMs
  • MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

    Wed Oct 7Imperial Ballroom #33All about evaluation
  • Legibility is Not Interpretability: Evaluating Judged Importance versus Actual Importance in Chain-Of-Thought Reasoning Steps

    Wed Oct 7Imperial Ballroom #34Science of LMs
  • Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

    Wed Oct 7Imperial Ballroom #35All about training
  • Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

    Wed Oct 7Imperial Ballroom #36All about training
  • Evolving Programmatic Skill Networks

    Wed Oct 7Imperial Ballroom #37LMs with tools and code
  • CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

    Wed Oct 7Imperial Ballroom #38LMs with tools and code
  • REVERE: Reflective Evolving Research Engineer

    Wed Oct 7Imperial Ballroom #39LMs with tools and code
  • p1: Better Prompt Optimization with Fewer Prompts

    Wed Oct 7Imperial Ballroom #40All about training
  • Compared to What? Baselines and Metrics for Counterfactual Prompting

    Wed Oct 7Imperial Ballroom #41All about evaluation
  • It’s How You Ask: Gender-Associated Linguistic Bias in LLMs

    Wed Oct 7Imperial Ballroom #42All about evaluation
  • Reach Into The Choir: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles

    Wed Oct 7Imperial Ballroom #43All about evaluation
  • Liberating LLM Capabilities in Full-Duplex Speech Models

    Wed Oct 7Imperial Ballroom #44Multimodal LMs
  • Learning to Draw ASCII Improves Spatial Reasoning in Language Models

    Wed Oct 7Imperial Ballroom #45Inference algorithms for LMs
  • Polar probe linearly decodes semantic structures from LLMs

    Wed Oct 7Imperial Ballroom #46Science of LMs
  • Large language models reorganize representational geometry during in-context learning

    Wed Oct 7Imperial Ballroom #47Science of LMs
  • In-Context Learning as Implicit Policy Gradient

    Wed Oct 7Imperial Ballroom #48Science of LMs
  • GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning

    Wed Oct 7Imperial Ballroom #49All about data
  • Risk Profiling and Modulation for LLMs

    Wed Oct 7Imperial Ballroom #50All about training
  • The Story Shapes the Agent: Narrative Priors in LLM Behavior

    Wed Oct 7Imperial Ballroom #51All about evaluation
  • Measuring Pragmatic Influence in LLM Instructions

    Wed Oct 7Imperial Ballroom #52All about evaluation
  • HieraSuite: A Holistic Toolkit for Building Versatile System-User Instruction Hierarchy

    Wed Oct 7Imperial Ballroom #53All about safety
  • LogicIF: Towards Complex Logic Instruction Following

    Wed Oct 7Imperial Ballroom #54All about training
  • Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis

    Wed Oct 7Imperial Ballroom #55LMs with tools and code
  • BugScope: Learn to Detect Bugs Like Human

    ORALWed Oct 7Imperial Ballroom #56LMs with tools and code
  • Coding Agents Don’t Know When to Act

    Wed Oct 7Imperial Ballroom #57LMs with tools and code
  • Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Wed Oct 7Imperial Ballroom #58All about evaluation
  • WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

    Wed Oct 7Imperial Ballroom #60Multimodal LMs
  • SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

    Wed Oct 7Imperial Ballroom #61All about evaluation
  • ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-art Discovery

    Wed Oct 7Imperial Ballroom #62LMs with tools and code
  • From Papers to Panoramas: Building Hierarchies of Scientific Literature at Scale

    Wed Oct 7Imperial Ballroom #63Diverse domains & novel applications
  • Document-as-Image Representations Fall Short for Scientific Retrieval

    Wed Oct 7Imperial Ballroom #64Multimodal LMs
  • Multimodal Latent Reasoning via Predictive Embeddings

    Wed Oct 7Imperial Ballroom #65Multimodal LMs
  • VERIFY A Benchmark of Visual Reasoning for Multimodal Reasoning Fidelity

    Wed Oct 7Imperial Ballroom #66Multimodal LMs
  • AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

    Wed Oct 7Imperial Ballroom #67Multimodal LMs
  • NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

    Wed Oct 7Imperial Ballroom #68Multimodal LMs
  • VisCodeBench: a new benchmark dataset for Text-to-Vis reflecting real-world practice

    Wed Oct 7Imperial Ballroom #69LMs with tools and code
  • CT-ΔBench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

    Wed Oct 7Imperial Ballroom #70Diverse domains & novel applications
  • CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction

    Wed Oct 7Imperial Ballroom #71Diverse domains & novel applications
  • SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching

    Wed Oct 7Imperial Ballroom #72Diverse domains & novel applications
  • TierCache: A Multi-Granularity Semantic Caching Framework for Structured Query Generation

    Wed Oct 7Imperial Ballroom #73Engineering for large LMs
  • Mitigating Knowledge Conflicts of Retrieval-Augmented Generation through Dual-Stage Confidence Measurement in Semantic Space

    Wed Oct 7Franciscan A #74LMs and the world
  • KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning

    Wed Oct 7Franciscan A #75LMs and the world
  • Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

    Wed Oct 7Franciscan A #76LMs and the world
  • Automatic Generation of High-Performance RL Environments

    Wed Oct 7Franciscan A #77Engineering for large LMs
  • What Makes a Sale? Simulating End-to-End Seller-Buyer Retail Dynamics with LLM Agents

    Wed Oct 7Franciscan A #78LMs and the world
  • CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V

    Wed Oct 7Franciscan A #79All about evaluation
  • VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents

    Wed Oct 7Franciscan A #80LMs with tools and code
  • iOSWorld: A Realistic iOS Environment for Benchmarking Phone Agents with Personalized User Identity and Memory

    Wed Oct 7Franciscan B #81LMs with tools and code
  • InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

    Wed Oct 7Franciscan B #82All about safety
  • InfMem: Learning System-2 Memory Control for Long-Context Agent

    Wed Oct 7Franciscan B #83LMs and the world
  • Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

    Wed Oct 7Franciscan B #85Compute-efficient LMs
  • ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering

    Wed Oct 7Franciscan B #86Compute-efficient LMs
  • RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning

    Wed Oct 7Franciscan B #87Compute-efficient LMs
  • Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking

    Wed Oct 7Franciscan B #88Diverse domains & novel applications
  • Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization

    Wed Oct 7Franciscan B #89LMs and interactions
  • CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

    Wed Oct 7Franciscan B #90All about safety
  • SOTOPIA-TOM: Evaluating Privacy and Information Management in Multi-Agent Interaction with Theory of Mind

    Wed Oct 7Franciscan C #91LMs and interactions
  • SNEAK: Evaluating Strategic Communication and Information Leakage in Large Language Models

    Wed Oct 7Franciscan C #92LMs and interactions
  • StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems

    Wed Oct 7Franciscan C #93LMs and interactions
  • Learning to Interrupt in Language-based Multi-agent Communication

    Wed Oct 7Franciscan C #94LMs and interactions
  • When Contextual Inference Fails: Cancelability in Interactive Instruction Following

    Wed Oct 7Franciscan C #95LMs and interactions
  • Commitment To Cooperation With Self-Negotiated Contracts

    Wed Oct 7Franciscan C #96LMs and interactions
  • Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harmful Behaviors

    Wed Oct 7Franciscan C #97All about safety
  • TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

    Wed Oct 7Franciscan C #98All about safety
  • Why Do Safety Guardrails Degrade Across Languages?

    Wed Oct 7Franciscan C #99All about safety
  • One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

    Wed Oct 7Franciscan C #100All about safety
  • Minimal, Local, Causal Explanations of Jailbreak Success in Large Language Models

    Wed Oct 7Grand Ballroom #101All about safety
  • Safety Cost of Steering Vectors Is Separable and Reducible

    Wed Oct 7Grand Ballroom #102All about safety
  • RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

    Wed Oct 7Grand Ballroom #103All about safety
  • Decoupled Alignment for Robust Plug-and-Play Adaptation

    Wed Oct 7Grand Ballroom #104All about safety
  • FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

    Wed Oct 7Grand Ballroom #105All about safety
  • Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

    Wed Oct 7Grand Ballroom #106All about safety
  • It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO

    Wed Oct 7Grand Ballroom #108All about safety
  • Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

    Wed Oct 7Grand Ballroom #109Multimodal LMs
  • Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

    Wed Oct 7Grand Ballroom #110Multimodal LMs
  • What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis

    Wed Oct 7Grand Ballroom #111Multimodal LMs
  • Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning

    Wed Oct 7Grand Ballroom #112All about training
  • Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

    ORALWed Oct 7Grand Ballroom #113All about training
  • On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning

    Wed Oct 7Grand Ballroom #114All about training
  • Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers

    Wed Oct 7Grand Ballroom #115Learning algorithms for LMs
  • How Transformers Learn to Plan via Multi-Token Prediction

    Wed Oct 7Grand Ballroom #116Science of LMs
  • MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers

    Wed Oct 7Grand Ballroom #117Science of LMs
  • Procedural Knowledge at Scale Improves Reasoning

    Wed Oct 7Grand Ballroom #118LMs and the world
  • QED-Nano: Teaching a Tiny Model to Prove Hard Theorems

    Wed Oct 7Grand Ballroom #119Diverse domains & novel applications
  • LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

    Wed Oct 7Grand Ballroom #120Diverse domains & novel applications
  • MathDuels: A Self-Play Benchmark That Grows

    Wed Oct 7Grand Ballroom #121Diverse domains & novel applications
  • BigCodeArena: A Platform for Executable Code Generation Evaluation

    Wed Oct 7Grand Ballroom #122LMs with tools and code
  • Meta-Harness: End-to-End Optimization of Model Harnesses

    Wed Oct 7Grand Ballroom #123LMs with tools and code
  • Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds

    Wed Oct 7Grand Ballroom #124All about safety
  • Start Classifying: Categorical Critics for LLM Reinforcement Learning

    Wed Oct 7Grand Ballroom #125All about training
  • A Rubric-Supervised Critic from Sparse Real-World Outcomes

    Wed Oct 7Grand Ballroom #126All about training
  • Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks

    Wed Oct 7Grand Ballroom #127All about training
  • Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

    Wed Oct 7Grand Ballroom #128All about training
  • Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing

    Wed Oct 7Grand Ballroom #129All about training
  • Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

    Wed Oct 7Grand Ballroom #130All about training
  • Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

    Wed Oct 7Grand Ballroom #131Multimodal LMs
  • OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

    Wed Oct 7Grand Ballroom #132Multimodal LMs
  • Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions

    Wed Oct 7Grand Ballroom #133Multimodal LMs
  • TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

    Wed Oct 7Grand Ballroom #134Multimodal LMs
  • FOCUS: Closed-Loop Attention Feedback for Efficient Vision-Language Understanding

    Wed Oct 7Grand Ballroom #135Compute-efficient LMs
  • Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

    Wed Oct 7Grand Ballroom #136Compute-efficient LMs
  • Multi-Granular Node Pruning for Causal Circuit Discovery

    Wed Oct 7Grand Ballroom #137Science of LMs
  • When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs

    Wed Oct 7Grand Ballroom #138Compute-efficient LMs
  • Lang-Prune: Unlocking Fair and Powerful Pruning for Multilingual Large Language Models

    Wed Oct 7Grand Ballroom #139Compute-efficient LMs
  • ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models

    Wed Oct 7Grand Ballroom #140Compute-efficient LMs
  • QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

    Wed Oct 7Grand Ballroom #141All about data
  • MidTool: Mid-training Data Synthesis for Agentic Tool Use

    Wed Oct 7Grand Ballroom #142All about data
  • Tool-Creating LLM Agents Gain Little from Keeping Their Tools

    Wed Oct 7Grand Ballroom #143LMs with tools and code
  • Do Humans and LLMs Diverge in Belief Revision? Evidence from a Bayesian Analysis

    ORALWed Oct 7Grand Ballroom #144Mind, brain, philosophy, law & LMs
  • Data-Efficient Adaptation of LLMs via Attention Head Reweighting

    Wed Oct 7Imperial Ballroom #1Compute-efficient LMs
  • A Prior-Aware Metric for Efficiently Distinguishing Memorization from Generalization in Large Language Models

    Wed Oct 7Imperial Ballroom #2Science of LMs
  • What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization

    Wed Oct 7Imperial Ballroom #3Science of LMs
  • Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

    ORALWed Oct 7Imperial Ballroom #4Science of LMs
  • An Investigation of Translationese in the Generations of Multilingual Large Language Models

    Wed Oct 7Imperial Ballroom #5LMs for everyone
  • Multilingual Agent-Based World Modeling for Social Science

    Wed Oct 7Imperial Ballroom #6LMs and the world
  • Social World Models

    Wed Oct 7Imperial Ballroom #7Mind, brain, philosophy, law & LMs
  • Neuro-Symbolic Synergy for World Modeling

    Wed Oct 7Imperial Ballroom #8LMs and the world
  • Neuro-symbolic PRM: Enhancing Scientific Reasoning via Structured Traces and Symbolic Verification

    Wed Oct 7Imperial Ballroom #9Inference algorithms for LMs
  • Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models

    Wed Oct 7Imperial Ballroom #10All about training
  • Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning

    Wed Oct 7Imperial Ballroom #11LMs and embodiment
  • Learning Next Action Predictors from Human-Computer Interaction

    Wed Oct 7Imperial Ballroom #12LMs and interactions
  • FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

    Wed Oct 7Imperial Ballroom #13Diverse domains & novel applications
  • YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution.

    Wed Oct 7Imperial Ballroom #14All about evaluation
  • DeliveryBench: Can Agents Earn Profit in Simulated Worlds?

    Wed Oct 7Imperial Ballroom #15LMs and the world
  • UrbanLLMind: Scalable LLM-Powered Urban Mobility Simulation with Open-Weight Models

    Wed Oct 7Imperial Ballroom #16LMs and the world
  • Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

    Wed Oct 7Imperial Ballroom #17LMs and the world
  • Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

    Wed Oct 7Imperial Ballroom #18LMs and interactions
  • Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry

    Wed Oct 7Imperial Ballroom #19LMs and interactions
  • The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

    Wed Oct 7Imperial Ballroom #20LMs and interactions
  • Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing

    Wed Oct 7Imperial Ballroom #21LMs and interactions
  • A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

    Wed Oct 7Imperial Ballroom #22Diverse domains & novel applications
  • PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing

    Wed Oct 7Imperial Ballroom #23Diverse domains & novel applications
  • OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

    Wed Oct 7Imperial Ballroom #24LMs with tools and code
  • MTA-Agent: An Open Recipe for Multimodal Deep Search Agents

    Wed Oct 7Imperial Ballroom #25LMs with tools and code
  • Bridging Databases and Documents: Data-Algorithm Co-Design for Hybrid Question Answering

    Wed Oct 7Imperial Ballroom #26LMs with tools and code
  • DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling

    Wed Oct 7Imperial Ballroom #27LMs with tools and code
  • CoreSemDB: Benchmarking Hybrid Semantic-Relational Query Processing over Text-Rich Databases

    Wed Oct 7Imperial Ballroom #28All about evaluation
  • VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean

    Wed Oct 7Imperial Ballroom #29LMs with tools and code
  • The Art of Building Verifiers for Computer Use Agents

    Wed Oct 7Imperial Ballroom #30LMs with tools and code
  • ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

    Wed Oct 7Imperial Ballroom #31LMs with tools and code
  • CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

    Wed Oct 7Imperial Ballroom #32LMs with tools and code
  • WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

    Wed Oct 7Imperial Ballroom #33LMs with tools and code
  • Back to the Future: A workbook time machine for spread- sheet creation benchmarks

    Wed Oct 7Imperial Ballroom #34All about evaluation
  • Many Ways to Be Fake: Benchmarking Fake News Detection Under Strategy-Driven AI Generation

    Wed Oct 7Imperial Ballroom #35All about safety
  • Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection

    Wed Oct 7Imperial Ballroom #36All about safety
  • How Humans and LLMs Read Gender into “Gender-Neutral” Physical Descriptions

    Wed Oct 7Imperial Ballroom #37LMs for everyone
  • From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models Through the lens of Social Relationship

    Wed Oct 7Imperial Ballroom #38LMs for everyone
  • Whose Standpoint do LLMs Reflect? Surfacing and Mitigating Epistemic Blindspots

    Wed Oct 7Imperial Ballroom #39LMs for everyone
  • Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

    Wed Oct 7Imperial Ballroom #40All about evaluation
  • Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

    Wed Oct 7Imperial Ballroom #41All about evaluation
  • Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling

    Wed Oct 7Imperial Ballroom #42Compute-efficient LMs
  • SCOPE: A Generative Approach for LLM Prompt Compression

    Wed Oct 7Imperial Ballroom #43Compute-efficient LMs
  • Optical Context Compression Is Just (Bad) Autoencoding

    Wed Oct 7Imperial Ballroom #44Compute-efficient LMs
  • SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Workloads

    Wed Oct 7Imperial Ballroom #45Compute-efficient LMs
  • Free(): Learning to Forget in Malloc-Only Reasoning Models

    Wed Oct 7Imperial Ballroom #46Compute-efficient LMs
  • Fork-think with Confidence

    Wed Oct 7Imperial Ballroom #47Inference algorithms for LMs
  • Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

    Wed Oct 7Imperial Ballroom #48Inference algorithms for LMs
  • Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs

    Wed Oct 7Imperial Ballroom #49Inference algorithms for LMs
  • From Mechanism Discovery to Proposal Closure: Graph-Grounded Hierarchical Search for Scientific Ideation

    Wed Oct 7Imperial Ballroom #50Diverse domains & novel applications
  • CrystalSTAR: Structured Action Orchestration with Trio-Reflection for Constrained Novel Crystal Discovery

    Wed Oct 7Imperial Ballroom #51Diverse domains & novel applications
  • MetaSymbO: Multi-Agent Language-Guided Metamaterial Discovery via Symbolic Latent Evolution

    Wed Oct 7Imperial Ballroom #52Diverse domains & novel applications
  • AMAT: Automated Multi-Agent Topology Design via Reinforcement Learning

    Wed Oct 7Imperial Ballroom #53LMs and interactions
  • Learning to Comply: Workflow-Grounded Environment Generation for Training Procedurally Compliant Agents

    Wed Oct 7Imperial Ballroom #54LMs with tools and code
  • RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

    Wed Oct 7Imperial Ballroom #55Multimodal LMs
  • PrefPO: Pairwise Preference Prompt Optimization

    Wed Oct 7Imperial Ballroom #56All about training
  • GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses

    Wed Oct 7Imperial Ballroom #57Diverse domains & novel applications
  • Process-Oriented Evaluation of AI-Assisted Scientific Writing

    Wed Oct 7Imperial Ballroom #58Diverse domains & novel applications
  • LLMs Corrupt Your Documents When You Delegate

    Wed Oct 7Imperial Ballroom #59LMs and interactions
  • Distributed Attacks in Persistent-State AI Control

    ORALWed Oct 7Imperial Ballroom #60All about safety
  • Preference Redirection via Attention Concentration: An Attack on Computer Use Agents

    Wed Oct 7Imperial Ballroom #61All about safety
  • Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

    Wed Oct 7Imperial Ballroom #62All about safety
  • FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

    Wed Oct 7Imperial Ballroom #63Multimodal LMs
  • RELISH: LLM REgression with a Latent Iterative State Head

    Wed Oct 7Imperial Ballroom #64Learning algorithms for LMs
  • A Comedy of Estimators: On KL Regularization in RL Training of LLMs

    Wed Oct 7Imperial Ballroom #65All about training
  • Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives

    Wed Oct 7Imperial Ballroom #66All about training
  • SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

    Wed Oct 7Imperial Ballroom #67Engineering for large LMs
  • NI Sampling++: Optimizing Token Order for Fast Discrete Diffusion Sampling with Reinforcement Learning

    Wed Oct 7Imperial Ballroom #68Inference algorithms for LMs
  • Mask-Aware Policy Gradients for Diffusion Language Models

    Wed Oct 7Imperial Ballroom #69All about training
  • Multi-Mask Diffusion Language Models for Few-Step Generation

    Wed Oct 7Imperial Ballroom #70Learning algorithms for LMs
  • Dual-Stream Decoding for Accelerated Large Language Models

    Wed Oct 7Imperial Ballroom #72Compute-efficient LMs
  • Trie Automata for Constrained Decoding over Large Finite Sets

    Wed Oct 7Imperial Ballroom #73Inference algorithms for LMs
  • When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

    Wed Oct 7Franciscan A #74Multimodal LMs
  • MURMUR: Cross-Lingual and Multimodal Retrieval-Augmented Reasoning for Open Question Answering in Tamil and Yoruba

    Wed Oct 7Franciscan A #75LMs for everyone
  • Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

    Wed Oct 7Franciscan A #76Multimodal LMs
  • Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts

    Wed Oct 7Franciscan A #77Multimodal LMs
  • The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

    Wed Oct 7Franciscan A #78Multimodal LMs
  • Seeing Isn’t Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

    Wed Oct 7Franciscan A #79Multimodal LMs
  • VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

    Wed Oct 7Franciscan A #80Multimodal LMs
  • Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

    Wed Oct 7Franciscan B #81Multimodal LMs
  • Multimodal Language Models Cannot Spot Spatial Inconsistencies

    Wed Oct 7Franciscan B #82Multimodal LMs
  • From Plausible to Grounded: Reinforcing Structured Consistency in Video MLLMs

    Wed Oct 7Franciscan B #83Multimodal LMs
  • TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs

    Wed Oct 7Franciscan B #84Multimodal LMs
  • Creativity and Hallucinations Share Mechanisms in LLMs

    Wed Oct 7Franciscan B #85Science of LMs
  • Disentangling MLP Neuron Weights in Vocabulary Space

    Wed Oct 7Franciscan B #86Science of LMs
  • Arithmetic in the Wild: Llama uses Standard Addition to Reason About Cyclic Concepts

    Wed Oct 7Franciscan B #87Science of LMs
  • Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs

    Wed Oct 7Franciscan B #88Mind, brain, philosophy, law & LMs
  • Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

    Wed Oct 7Franciscan B #89Science of LMs
  • LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

    Wed Oct 7Franciscan B #90Science of LMs
  • Evaluating steering techniques using human similarity judgments

    Wed Oct 7Franciscan C #91Mind, brain, philosophy, law & LMs
  • Steering Awareness: Detecting Activation Steering from Within

    Wed Oct 7Franciscan C #92All about safety
  • Steering Instruction Hierarchies at Inference Time

    Wed Oct 7Franciscan C #93All about safety
  • Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

    Wed Oct 7Franciscan C #94LMs and interactions
  • SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

    Wed Oct 7Franciscan C #95LMs and interactions
  • COMPASS: Benchmarking Constrained Optimization in LLM Agents

    Wed Oct 7Franciscan C #96LMs with tools and code
  • Analysis of Optimality of Large Language Models on Planning Problems

    Wed Oct 7Franciscan C #97Inference algorithms for LMs
  • The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning

    Wed Oct 7Franciscan C #98Inference algorithms for LMs
  • KTPO: K-Step Test-Time Policy Optimization for Long- Horizon Discovery

    Wed Oct 7Franciscan C #99Inference algorithms for LMs
  • ACTOR-CURATOR: Online Curriculum Learning via Policy Improvement Bandits for RL Post-Training

    Wed Oct 7Franciscan C #100All about training
  • Self-Evolving Curriculum for LLM Reasoning

    Wed Oct 7Grand Ballroom #101All about training
  • New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR

    Wed Oct 7Grand Ballroom #102Science of LMs
  • Beyond Distribution Sharpening: The Importance of Task Rewards

    ORALWed Oct 7Grand Ballroom #103All about training
  • From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

    Wed Oct 7Grand Ballroom #104All about training
  • Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs

    Wed Oct 7Grand Ballroom #105Diverse domains & novel applications
  • CONCORD: Label-Free Calibration of Verbalized LLM Confidence via Rollout Consistency

    Wed Oct 7Grand Ballroom #106Science of LMs
  • PAM: Training Moderation Filters from Policy-Derived Supervision

    Wed Oct 7Grand Ballroom #107All about safety
  • Quality of Rejections Matters: Preference Data Construction for Faithful Summarization

    Wed Oct 7Grand Ballroom #108All about data
  • Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling

    Wed Oct 7Grand Ballroom #109All about evaluation
  • From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

    Wed Oct 7Grand Ballroom #110All about evaluation
  • SWE-chat: Coding Agent Interactions From Real Users in the Wild

    Wed Oct 7Grand Ballroom #111LMs with tools and code
  • The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior

    Wed Oct 7Grand Ballroom #112LMs with tools and code
  • FitText: Evolving Agent Tool Ecologies via Memetic Retrieval

    Wed Oct 7Grand Ballroom #113LMs with tools and code
  • CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

    Wed Oct 7Grand Ballroom #114LMs and the world
  • When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies

    Wed Oct 7Grand Ballroom #115LMs and the world
  • Only Ask What You Don’t Know: Grounded Delta Planning for Efficient Multi-step RAG

    Wed Oct 7Grand Ballroom #116LMs and the world
  • RAQE: Reranker-Aligned Query Expansion via Label-Free Group-Relative Policy Optimization

    Wed Oct 7Grand Ballroom #117LMs and the world
  • AgentIR: Reasoning-Aware Retrieval for Deep Research Agents

    Wed Oct 7Grand Ballroom #118LMs with tools and code
  • ResearcherBench: Evaluating Deep AI Research Systems on Open-ended AI Research Tasks

    Wed Oct 7Grand Ballroom #119All about evaluation
  • Sci-VBench: Evaluating Knowledge- and Reasoning- Intensive Video Generation in Science Domains

    Wed Oct 7Grand Ballroom #120Multimodal LMs
  • ScienceMeter: Tracking Scientific Knowledge Updates in Language Models

    Wed Oct 7Grand Ballroom #121LMs and the world
  • The Shrinking Lifespan of LLMs in Science

    Wed Oct 7Grand Ballroom #122Diverse domains & novel applications
  • The Hidden Cost of Thinking: Energy Use and Environmental Impact of LMs Beyond Pretraining

    Wed Oct 7Grand Ballroom #123Engineering for large LMs
  • Data-efficient pre-training by scaling synthetic megadocs

    Wed Oct 7Grand Ballroom #124All about data
  • Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

    Wed Oct 7Grand Ballroom #125All about data
  • How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

    ORALWed Oct 7Grand Ballroom #126All about data
  • FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale

    Wed Oct 7Grand Ballroom #127All about data
  • LM-mixup: Text Data Augmentation via Language Model based Mixup

    Wed Oct 7Grand Ballroom #128All about data
  • Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection

    Wed Oct 7Grand Ballroom #129All about training
  • Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans

    Wed Oct 7Grand Ballroom #130Mind, brain, philosophy, law & LMs
  • In-Context Examples Suppress Scientific Knowledge Recall in LLMs

    Wed Oct 7Grand Ballroom #131All about training
  • Improving Latent Generalization Using Test-time Compute

    Wed Oct 7Grand Ballroom #132Inference algorithms for LMs
  • Language Models That Think, Chat Better

    Wed Oct 7Grand Ballroom #133All about training
  • Reasoning on a Spectrum: Aligning LLMs to System 1 and System 2 Thinking

    Wed Oct 7Grand Ballroom #134Mind, brain, philosophy, law & LMs
  • Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond

    Wed Oct 7Grand Ballroom #135Learning algorithms for LMs
  • Are Latent Reasoning Models Easily Interpretable?

    Wed Oct 7Grand Ballroom #136Science of LMs
  • Grounding latent algorithm routing in transformer reasoning

    Wed Oct 7Grand Ballroom #137Science of LMs
  • LLM Router: Rethinking Routing with Prefill Activations

    Wed Oct 7Grand Ballroom #138Inference algorithms for LMs
  • No Single Best Model for Diversity: Learning a Router for Sample Diversity

    Wed Oct 7Grand Ballroom #139Inference algorithms for LMs
  • Routing Entropy: A Hidden Self-Verifier for Free in Mixture-of-Experts LLMs

    Wed Oct 7Grand Ballroom #140Learning algorithms for LMs
  • Reinforcement Routing for Mixtures of LoRAs in Parameter-Efficient LLM Finetuning

    Wed Oct 7Grand Ballroom #141Compute-efficient LMs
  • ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale

    Wed Oct 7Grand Ballroom #142Engineering for large LMs
  • Let’s Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

    Wed Oct 7Grand Ballroom #143Engineering for large LMs
  • Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?

    Thu Oct 8Imperial Ballroom #1All about safety
  • The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs

    Thu Oct 8Imperial Ballroom #2All about safety
  • Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

    Thu Oct 8Imperial Ballroom #3All about safety
  • Efficient Safety Alignment of Language Models via Latent Personality Traits

    Thu Oct 8Imperial Ballroom #4All about safety
  • Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

    Thu Oct 8Imperial Ballroom #5All about safety
  • From Geometry to Behavior: How Instruction Exposure Unlocks Latent Rhetorical Directions in LLMs

    Thu Oct 8Imperial Ballroom #6Science of LMs
  • Latent Structure of Affective Representations in Large Language Models

    Thu Oct 8Imperial Ballroom #7Science of LMs
  • Attractor States Emerge in Multi-Turn LLM Conversations

    Thu Oct 8Imperial Ballroom #8LMs and interactions
  • STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts

    Thu Oct 8Imperial Ballroom #9Inference algorithms for LMs
  • Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

    Thu Oct 8Imperial Ballroom #10Mind, brain, philosophy, law & LMs
  • DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models

    Thu Oct 8Imperial Ballroom #11LMs and embodiment
  • Decoupling Planning and Control for Instructable Agents

    Thu Oct 8Imperial Ballroom #12LMs and embodiment
  • Aligning Language Models from User Interactions

    Thu Oct 8Imperial Ballroom #13LMs and interactions
  • Peer-Predictive Self-Training for Language Model Reasoning

    Thu Oct 8Imperial Ballroom #14All about training
  • Variational Co-Evolution via Reinforcement Learning

    Thu Oct 8Imperial Ballroom #15All about training
  • Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

    Thu Oct 8Imperial Ballroom #16All about training
  • On Epistemic Diversity in Large Language Models

    Thu Oct 8Imperial Ballroom #17All about evaluation
  • CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse

    Thu Oct 8Imperial Ballroom #18Science of LMs
  • Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation

    Thu Oct 8Imperial Ballroom #19Learning algorithms for LMs
  • Reasoning Fine-Tuning Induces Persistent Latent Policy States

    Thu Oct 8Imperial Ballroom #20Science of LMs
  • Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing

    Thu Oct 8Imperial Ballroom #21Science of LMs
  • A Mirage of Coherence: How Metaphor Impacts Language Models' Discourse Coherence Assessment

    ORALThu Oct 8Imperial Ballroom #22Mind, brain, philosophy, law & LMs
  • Should We be Pedantic About Reasoning Errors in Machine Translation?

    Thu Oct 8Imperial Ballroom #23LMs for everyone
  • MELD: Multilingual Ensemble via Logical Debate

    Thu Oct 8Imperial Ballroom #24LMs for everyone
  • TowerVision: Understanding and Improving Multilinguality in Vision-Language Models

    Thu Oct 8Imperial Ballroom #25LMs for everyone
  • LF²AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

    Thu Oct 8Imperial Ballroom #26Multimodal LMs
  • DARE: Diffusion Large Language Models Alignment and Reinforcement Executor

    Thu Oct 8Imperial Ballroom #27All about training
  • A Tale of Two Temperatures: Simple, Efficient, and Diverse Sampling from Diffusion Language Models

    Thu Oct 8Imperial Ballroom #28Inference algorithms for LMs
  • Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

    Thu Oct 8Imperial Ballroom #29Science of LMs
  • Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning

    Thu Oct 8Imperial Ballroom #30All about training
  • Uniform Information Density-Based Preference Optimization

    Thu Oct 8Imperial Ballroom #31All about training
  • A11yn: Aligning LLMs for Web Accessibility-Aware UI Generation

    Thu Oct 8Imperial Ballroom #32LMs with tools and code
  • Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?

    Thu Oct 8Imperial Ballroom #33Multimodal LMs
  • GLAZE: A Gaze-Language Benchmark for Grounding Human Gaze on Web User Interfaces

    Thu Oct 8Imperial Ballroom #34Multimodal LMs
  • Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

    Thu Oct 8Imperial Ballroom #35Multimodal LMs
  • BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

    Thu Oct 8Imperial Ballroom #36All about evaluation
  • Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

    Thu Oct 8Imperial Ballroom #37All about evaluation
  • Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

    Thu Oct 8Imperial Ballroom #38All about evaluation
  • Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

    Thu Oct 8Imperial Ballroom #39All about evaluation
  • Comparing Developer and LLM Biases in Code Evaluation

    Thu Oct 8Imperial Ballroom #40All about evaluation
  • Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

    Thu Oct 8Imperial Ballroom #41All about evaluation
  • Agent Psychometrics: Task-Level Performance Prediction in Agentic Coding Benchmarks

    Thu Oct 8Imperial Ballroom #42All about evaluation
  • SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

    Thu Oct 8Imperial Ballroom #43LMs with tools and code
  • Constraint decay: The Fragility of LLM Agents in Backend Code Generation

    Thu Oct 8Imperial Ballroom #44LMs with tools and code
  • When Correct Patches Become Vulnerable: Red Teaming LLM-Based Program Repair Agents

    Thu Oct 8Imperial Ballroom #45All about safety
  • CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering

    Thu Oct 8Imperial Ballroom #46All about safety
  • FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?

    Thu Oct 8Imperial Ballroom #47Diverse domains & novel applications
  • AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems

    Thu Oct 8Imperial Ballroom #48Diverse domains & novel applications
  • Routesplain: Towards Faithful and Intervenable Routing for Software-related Tasks

    Thu Oct 8Imperial Ballroom #49LMs with tools and code
  • SkillRouter: Skill Routing for LLM Agents at Scale

    Thu Oct 8Imperial Ballroom #50LMs with tools and code
  • How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

    Thu Oct 8Imperial Ballroom #51LMs with tools and code
  • EvoSkill: Automated Skill Discovery for Multi-Agent Systems

    Thu Oct 8Imperial Ballroom #52LMs with tools and code
  • FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

    Thu Oct 8Imperial Ballroom #53LMs with tools and code
  • Dr. Zero: Self-Evolving Search Agents without Training Data

    Thu Oct 8Imperial Ballroom #54LMs with tools and code
  • AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback

    Thu Oct 8Imperial Ballroom #55All about training
  • Neural Architecture Discovery via Autonomous Evolution

    Thu Oct 8Imperial Ballroom #56Learning algorithms for LMs
  • CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

    Thu Oct 8Imperial Ballroom #57LMs and interactions
  • EvoX: Meta-Evolution for Automated Discovery

    Thu Oct 8Imperial Ballroom #58Learning algorithms for LMs
  • EvoLen: Evolution-Guided Tokenization for DNA Language Model

    Thu Oct 8Imperial Ballroom #59Diverse domains & novel applications
  • Struc-EMB: The Potential of Structure-Aware Encoding in Language Embeddings

    Thu Oct 8Imperial Ballroom #60Learning algorithms for LMs
  • Structure Before Collapse: Transient Semantic Geometry in Next-Token Prediction

    Thu Oct 8Imperial Ballroom #61Science of LMs
  • Why Your Prompt Compressor Is Deleting the Wrong Tokens: An Information-Theoretic Diagnosis

    Thu Oct 8Imperial Ballroom #62Compute-efficient LMs
  • Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

    Thu Oct 8Imperial Ballroom #63Compute-efficient LMs
  • CeRA: Breaking the Linear Ceiling of Low-Rank Adaptation with Non-linearity Retained at Inference

    Thu Oct 8Imperial Ballroom #64Compute-efficient LMs
  • Super Weights in LLMs and the Failure of Selective Training

    Thu Oct 8Imperial Ballroom #65Science of LMs
  • Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning

    Thu Oct 8Imperial Ballroom #66Learning algorithms for LMs
  • SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

    Thu Oct 8Imperial Ballroom #67Compute-efficient LMs
  • Direct Multi-Token Decoding

    Thu Oct 8Imperial Ballroom #68Compute-efficient LMs
  • Accelerating Speculative Decoding with Block Diffusion Draft Trees

    Thu Oct 8Imperial Ballroom #69Compute-efficient LMs
  • FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents

    Thu Oct 8Imperial Ballroom #70LMs with tools and code
  • BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

    Thu Oct 8Imperial Ballroom #71Diverse domains & novel applications
  • DeepScholar-Bench: A Live Benchmark for Automated Evaluation of Generative Research Synthesis

    Thu Oct 8Imperial Ballroom #72All about evaluation
  • WebReal: Benchmarking DeepSearch Agents on Real-world User Information Needs

    Thu Oct 8Imperial Ballroom #73All about evaluation
  • Towards Comprehensive Mobile Application Testing for User-centric Feature Coverage via Simulating Real-World Users with Personas

    Thu Oct 8Franciscan A #74LMs with tools and code
  • CocoaBench: Evaluating unified digital agents in the wild

    Thu Oct 8Franciscan A #75All about evaluation
  • AlphaEval: Evaluating Agents in Production

    Thu Oct 8Franciscan A #76All about evaluation
  • ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

    Thu Oct 8Franciscan A #77All about safety
  • ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

    Thu Oct 8Franciscan A #78All about safety
  • PM-Bench: Evaluating Prospective Memory of LLM Agents

    Thu Oct 8Franciscan A #79All about evaluation
  • AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

    Thu Oct 8Franciscan A #80LMs and interactions
  • MT-PINGEVAL: Evaluating Multi-Turn Collaboration with Private Information Games

    Thu Oct 8Franciscan B #81LMs and interactions
  • AI Assistance Reduces Persistence and Hurts Independent Performance

    Thu Oct 8Franciscan B #82LMs and interactions
  • Verbalizing LLMs' assumptions to explain and control sycophancy

    Thu Oct 8Franciscan B #83All about safety
  • Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks

    Thu Oct 8Franciscan B #84All about evaluation
  • Synthetic Sandbox for Training ML Engineering Agents

    Thu Oct 8Franciscan B #85All about training
  • TritonRL: Training LLMs to Think and Code Triton Without Cheating

    Thu Oct 8Franciscan B #86LMs with tools and code
  • Quasar: A Programming Language Specialized for LLM Code Actions

    Thu Oct 8Franciscan B #87LMs with tools and code
  • MetaLint: Easy-to-Hard Generalization for Code Linting

    Thu Oct 8Franciscan B #88LMs with tools and code
  • MAPLE: Metadata Augmented Private Language Evolution

    Thu Oct 8Franciscan B #89All about data
  • PolicyLong: Towards On-Policy Context Extension

    Thu Oct 8Franciscan B #90All about data
  • Reflective Context Learning: Studying the Optimization Primitives of Context Space

    Thu Oct 8Franciscan C #91All about training
  • Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover

    Thu Oct 8Franciscan C #92Learning algorithms for LMs
  • Recovering Wasted Compute in Autoresearch Agents

    Thu Oct 8Franciscan C #93Compute-efficient LMs
  • Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection

    Thu Oct 8Franciscan C #94Science of LMs
  • Prescriptive Scaling Laws for Data Constrained Training

    Thu Oct 8Franciscan C #95Science of LMs
  • The Finetuner’s Fallacy: When to Pretrain with Your Finetuning Data

    Thu Oct 8Franciscan C #96All about data
  • Domain-Aware Scaling Laws Uncover Data Synergy

    Thu Oct 8Franciscan C #97Science of LMs
  • Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning

    Thu Oct 8Franciscan C #98Science of LMs
  • Do Language Models Consistently Encode the Current Year?

    Thu Oct 8Franciscan C #99Science of LMs
  • Studying the Soupability of Documents in State Space Models

    Thu Oct 8Franciscan C #100Learning algorithms for LMs
  • Message Passing Enables Efficient Reasoning

    ORALThu Oct 8Grand Ballroom #101Compute-efficient LMs
  • Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence

    Thu Oct 8Grand Ballroom #102Learning algorithms for LMs
  • Disentangling the Expressivity of RoPE

    Thu Oct 8Grand Ballroom #103Learning algorithms for LMs
  • Compliments to the Complementizer: Indirect Routing of Syntactic Structure in a Transformer Language Model

    Thu Oct 8Grand Ballroom #104Science of LMs
  • Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge

    Thu Oct 8Grand Ballroom #105Science of LMs
  • KAMR: Grounding Generation via Knowledge-Aligned Multi-hop Retrieval

    Thu Oct 8Grand Ballroom #106LMs and the world
  • CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning

    Thu Oct 8Grand Ballroom #107LMs and the world
  • SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

    Thu Oct 8Grand Ballroom #108LMs and the world
  • Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Differentially Private

    Thu Oct 8Grand Ballroom #109All about safety
  • Evaluating the Retrieval Robustness of Large Language Models

    Thu Oct 8Grand Ballroom #110All about evaluation
  • Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

    Thu Oct 8Grand Ballroom #111All about evaluation
  • QualAlign: Benchmarking Automated Qualitative Coding Against Human Schemas

    Thu Oct 8Grand Ballroom #112All about evaluation
  • How social context shapes value-related content in Large Language Model outputs

    Thu Oct 8Grand Ballroom #113Mind, brain, philosophy, law & LMs
  • Ideological Bias in LLMs’ Economic Causal Reasoning

    Thu Oct 8Grand Ballroom #114LMs and the world
  • Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations

    Thu Oct 8Grand Ballroom #115LMs and the world
  • Incentive-Aware Multi-Fidelity Optimization for Generative Advertising in Large Language Models

    Thu Oct 8Grand Ballroom #116LMs and the world
  • SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation

    Thu Oct 8Grand Ballroom #117All about training
  • Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards

    Thu Oct 8Grand Ballroom #118All about training
  • CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

    Thu Oct 8Grand Ballroom #119Multimodal LMs
  • Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation

    Thu Oct 8Grand Ballroom #120Diverse domains & novel applications
  • What do your logits know?

    Thu Oct 8Grand Ballroom #121Science of LMs
  • Attr-Kit: An Efficient Toolkit for No-Decode Source Attribution

    Thu Oct 8Grand Ballroom #122Science of LMs
  • Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?

    Thu Oct 8Grand Ballroom #123Science of LMs
  • When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

    Thu Oct 8Grand Ballroom #124Science of LMs
  • Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

    Thu Oct 8Grand Ballroom #125Learning algorithms for LMs
  • Extracting memorized pieces of (copyrighted) books from open-weight language models

    ORALThu Oct 8Grand Ballroom #126All about safety
  • Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

    Thu Oct 8Grand Ballroom #127All about safety
  • Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

    Thu Oct 8Grand Ballroom #128All about safety
  • Detecting Safety Violations Across Many Agent Traces

    Thu Oct 8Grand Ballroom #129All about safety
  • Training Agents to Self-Report Misbehavior

    Thu Oct 8Grand Ballroom #130All about safety
  • Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents

    Thu Oct 8Grand Ballroom #131All about safety
  • BrowseSafe: Understanding and Detecting Prompt Injection Within AI Browser Agents

    Thu Oct 8Grand Ballroom #132All about safety
  • CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

    Thu Oct 8Grand Ballroom #133All about safety
  • Extended to Reality: Prompt Injection in 3D Environments

    Thu Oct 8Grand Ballroom #134All about safety
  • Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

    Thu Oct 8Grand Ballroom #135All about safety
  • When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets

    Thu Oct 8Grand Ballroom #136All about evaluation
  • TRUST: A Crowdsourcing Framework for Auditing Large Language Model Reasoning

    Thu Oct 8Grand Ballroom #137All about evaluation
  • Auditing Moderation Robustness Under Realistic Settings

    Thu Oct 8Grand Ballroom #138All about safety
  • Back to Basics: Revisiting ASR in the Age of Voice Agents

    Thu Oct 8Grand Ballroom #139Multimodal LMs
  • SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers

    Thu Oct 8Grand Ballroom #140LMs and interactions
  • Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

    Thu Oct 8Grand Ballroom #141All about safety
  • Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules

    Thu Oct 8Grand Ballroom #142All about safety
  • When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

    Thu Oct 8Grand Ballroom #143All about training
  • Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

    ORALThu Oct 8Grand Ballroom #144All about safety
  • Lost in Distraction: LLM Planning Agents Recognize but Fail to Utilize Relevant Information under Context Noise

    Thu Oct 8Imperial Ballroom #1Inference algorithms for LMs
  • Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find

    Thu Oct 8Imperial Ballroom #2All about evaluation
  • Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

    Thu Oct 8Imperial Ballroom #3Learning algorithms for LMs
  • Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths

    Thu Oct 8Imperial Ballroom #4Learning algorithms for LMs
  • MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution

    Thu Oct 8Imperial Ballroom #5Learning algorithms for LMs
  • DepthSSD: Rethinking Residual Connections via State Space Models on the Depth Axis

    Thu Oct 8Imperial Ballroom #6Learning algorithms for LMs
  • MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference

    Thu Oct 8Imperial Ballroom #7Compute-efficient LMs
  • ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

    Thu Oct 8Imperial Ballroom #8Inference algorithms for LMs
  • SPEED: Specialized Position Experts for Efficient Speculative Decoding

    Thu Oct 8Imperial Ballroom #9Inference algorithms for LMs
  • Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding

    Thu Oct 8Imperial Ballroom #10Inference algorithms for LMs
  • Statistically-Lossless Quantization of Large Language Models

    Thu Oct 8Imperial Ballroom #11Compute-efficient LMs
  • NIRVANA: Structured Pruning Reimagined for Large Language Model Compression

    Thu Oct 8Imperial Ballroom #12Compute-efficient LMs
  • OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation

    Thu Oct 8Imperial Ballroom #13Compute-efficient LMs
  • Principled and Scalable Diversity-Aware Retrieval via Cardinality-Constrained Binary Quadratic Programming

    Thu Oct 8Imperial Ballroom #14LMs and the world
  • Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision–Language Models

    Thu Oct 8Imperial Ballroom #15All about data
  • Opusanimation: Code-based dynamic chart generation

    Thu Oct 8Imperial Ballroom #16Multimodal LMs
  • Co-Director: Agentic Generative Video Storytelling

    Thu Oct 8Imperial Ballroom #17Multimodal LMs
  • Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs

    Thu Oct 8Imperial Ballroom #18Multimodal LMs
  • Detecting and Suppressing Reward Hacking with Gradient Fingerprints

    Thu Oct 8Imperial Ballroom #19All about training
  • Voice of Reason: Reinforcement Learning for Spoken Math

    Thu Oct 8Imperial Ballroom #20Multimodal LMs
  • Do We Need Frontier Models to Verify Mathematical Proofs?

    Thu Oct 8Imperial Ballroom #21Inference algorithms for LMs
  • Ill-Defined Math: Benchmarking LLM Reasoning Beyond Well-Defined Problems

    Thu Oct 8Imperial Ballroom #22All about evaluation
  • Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

    Thu Oct 8Imperial Ballroom #23All about training
  • GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics

    Thu Oct 8Imperial Ballroom #24Science of LMs
  • Grammatical ``grandmother neurons'' are rare in LLMs

    Thu Oct 8Imperial Ballroom #25Science of LMs
  • Vocabulary embeddings organize linguistic structure early in language model training

    Thu Oct 8Imperial Ballroom #26Science of LMs
  • Lost in Backpropagation: The LM Head is a Gradient Bottleneck

    Thu Oct 8Imperial Ballroom #27Engineering for large LMs
  • BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs

    Thu Oct 8Imperial Ballroom #28Multimodal LMs
  • REFRAMED: Towards Realistic Audio Description Generation for Movies

    Thu Oct 8Imperial Ballroom #29Multimodal LMs
  • TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics

    Thu Oct 8Imperial Ballroom #30LMs and interactions
  • MoltNet: Understanding Social Behavior of AI Agents in the Agent-Native MoltBook

    Thu Oct 8Imperial Ballroom #31LMs and interactions
  • CooperBench: Why coding agents cannot be your teammates yet

    Thu Oct 8Imperial Ballroom #32LMs and interactions
  • When Does Personality Composition Matter for Multi-Agent LLM Teams?

    Thu Oct 8Imperial Ballroom #33LMs and interactions
  • High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination

    Thu Oct 8Imperial Ballroom #34LMs and interactions
  • Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

    Thu Oct 8Imperial Ballroom #35Mind, brain, philosophy, law & LMs
  • Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest

    Thu Oct 8Imperial Ballroom #36Mind, brain, philosophy, law & LMs
  • “I Didn’t Make the Micro Decisions”: Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration

    Thu Oct 8Imperial Ballroom #37LMs and interactions
  • CCBench: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

    Thu Oct 8Imperial Ballroom #38LMs for everyone
  • In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

    Thu Oct 8Imperial Ballroom #39LMs for everyone
  • RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

    Thu Oct 8Imperial Ballroom #40All about evaluation
  • ConsumerBench: Benchmarking Generative AI Applications on End-User Devices

    Thu Oct 8Imperial Ballroom #41Engineering for large LMs
  • MolDesignBench: Evaluating LLM-based Agent for Scenario- grounded Molecular Design

    Thu Oct 8Imperial Ballroom #42Diverse domains & novel applications
  • RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

    Thu Oct 8Imperial Ballroom #43Diverse domains & novel applications
  • Can Coding Agents Reproduce Findings in Computational Materials Science?

    Thu Oct 8Imperial Ballroom #44Diverse domains & novel applications
  • Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints

    Thu Oct 8Imperial Ballroom #45Diverse domains & novel applications
  • Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Thu Oct 8Imperial Ballroom #46LMs with tools and code
  • Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

    Thu Oct 8Imperial Ballroom #47All about evaluation
  • Agent Alpha: Unifying Generation, Exploration and Evaluation via Tree Search for Computer-Use Solution Discovery

    Thu Oct 8Imperial Ballroom #48Inference algorithms for LMs
  • Meta-Reinforcement Learning with Self-Reflection for Agentic Search

    Thu Oct 8Imperial Ballroom #49All about training
  • Cycle-Consistent Search: Question Reconstructability as a Proxy Reward for Search Agent Training

    Thu Oct 8Imperial Ballroom #50All about training
  • OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data

    Thu Oct 8Imperial Ballroom #51All about data
  • GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum

    Thu Oct 8Imperial Ballroom #52LMs and the world
  • Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge

    Thu Oct 8Imperial Ballroom #53LMs with tools and code
  • TierMem: Balancing Compressed Memory and Raw Evidence for Long-Horizon Agent Memory

    Thu Oct 8Imperial Ballroom #54LMs with tools and code
  • Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression

    Thu Oct 8Imperial Ballroom #55Compute-efficient LMs
  • Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought

    Thu Oct 8Imperial Ballroom #56Compute-efficient LMs
  • Fractured Chain-of-Thought Reasoning

    Thu Oct 8Imperial Ballroom #57Inference algorithms for LMs
  • Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

    Thu Oct 8Imperial Ballroom #58All about training
  • Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning

    Thu Oct 8Imperial Ballroom #59Inference algorithms for LMs
  • Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

    Thu Oct 8Imperial Ballroom #60All about evaluation
  • TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

    Thu Oct 8Imperial Ballroom #61Multimodal LMs
  • Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

    Thu Oct 8Imperial Ballroom #62Multimodal LMs
  • OpenStamp: A Watermark for Open-Source Language Models

    Thu Oct 8Imperial Ballroom #63All about safety
  • Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs

    Thu Oct 8Imperial Ballroom #64All about safety
  • Human vs Machine Translation Detection: A Cross-Model, Cross-Domain, and Low-Resource Analysis

    Thu Oct 8Imperial Ballroom #65LMs for everyone
  • We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference

    Thu Oct 8Imperial Ballroom #66Mind, brain, philosophy, law & LMs
  • PromptNCE: Conditional Probabilities and PMI Using Only LLMs and Contrastive Estimation Prompts

    Thu Oct 8Imperial Ballroom #67Science of LMs
  • Understanding In-context Learning of Addition via Activation Subspaces

    Thu Oct 8Imperial Ballroom #69Science of LMs
  • Reasoning Models Know What’s Important, and Encode It in Their Activations

    Thu Oct 8Imperial Ballroom #70Science of LMs
  • Component and Dimension Sparsity in Transformer Refusal Mechanisms

    Thu Oct 8Imperial Ballroom #71Science of LMs
  • Prism-Δ: Differential Subspace Steering for Prompt Highlighting in Large Language Models

    Thu Oct 8Imperial Ballroom #72Inference algorithms for LMs
  • Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization

    Thu Oct 8Imperial Ballroom #73Learning algorithms for LMs
  • Model Merging via Data-Free Covariance Estimation

    Thu Oct 8Franciscan A #74Learning algorithms for LMs
  • Multi-objective Evolutionary Merging Enables Efficient Reasoning Models

    Thu Oct 8Franciscan A #75Compute-efficient LMs
  • ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization

    Thu Oct 8Franciscan A #76Diverse domains & novel applications
  • Learning Reusable Program Transformations via LLM-Guided Rule Synthesis

    Thu Oct 8Franciscan A #77LMs with tools and code
  • Models Can Model, But Can’t Bind: Structured Grounding in Text-to-Optimization

    Thu Oct 8Franciscan A #78Diverse domains & novel applications
  • The Format Tax: Separating the Cost of Requesting Structure from the Cost of Enforcing It

    Thu Oct 8Franciscan A #79Inference algorithms for LMs
  • Syntax Without Semantics: Teaching LLMs to Code in an Unseen Language

    Thu Oct 8Franciscan A #80LMs with tools and code
  • Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization

    Thu Oct 8Franciscan B #81LMs with tools and code
  • Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts

    Thu Oct 8Franciscan B #82Learning algorithms for LMs
  • DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models

    Thu Oct 8Franciscan B #83All about training
  • MoANT: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning

    Thu Oct 8Franciscan B #84All about training
  • MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

    Thu Oct 8Franciscan B #85LMs for everyone
  • Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

    Thu Oct 8Franciscan B #86LMs for everyone
  • PreMoE: Proactive Inference for Efficient Mixture-of-Experts

    Thu Oct 8Franciscan B #87Compute-efficient LMs
  • Path-Constrained Mixture-of-Experts

    Thu Oct 8Franciscan B #88Learning algorithms for LMs
  • Expert Threshold Routing for Autoregressive Language Modeling with Dynamic Computation Allocation and Load Balancing

    Thu Oct 8Franciscan B #89Learning algorithms for LMs
  • Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models

    Thu Oct 8Franciscan B #90Learning algorithms for LMs
  • TRIMS: Trajectory-Ranked Instruction Masked Supervision for Diffusion Language Models

    Thu Oct 8Franciscan C #91Learning algorithms for LMs
  • Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

    Thu Oct 8Franciscan C #92Inference algorithms for LMs
  • MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

    Thu Oct 8Franciscan C #93Learning algorithms for LMs
  • CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

    Thu Oct 8Franciscan C #94Inference algorithms for LMs
  • RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback

    Thu Oct 8Franciscan C #95All about training
  • Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

    Thu Oct 8Franciscan C #96All about evaluation
  • Mitigating LLM biases toward spurious social contexts using direct preference optimization

    Thu Oct 8Franciscan C #97LMs for everyone
  • WorldPM: Scaling Human Preference Modeling via Real-World Feedback

    Thu Oct 8Franciscan C #98All about training
  • Beyond expert users: agents should help users construct preferences, not just elicit them

    Thu Oct 8Franciscan C #99LMs and interactions
  • Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation

    Thu Oct 8Franciscan C #100LMs and interactions
  • LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers

    Thu Oct 8Grand Ballroom #101Science of LMs
  • Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting

    Thu Oct 8Grand Ballroom #102Mind, brain, philosophy, law & LMs
  • INSIDE the Student’s Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

    Thu Oct 8Grand Ballroom #103Diverse domains & novel applications
  • HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning

    Thu Oct 8Grand Ballroom #104Inference algorithms for LMs
  • Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems

    Thu Oct 8Grand Ballroom #105LMs and interactions
  • Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning

    Thu Oct 8Grand Ballroom #106LMs and interactions
  • Stabilizing Off-Policy Training for Long-Horizon LLM Agents via Turn-Level Importance Sampling and Clipping-Triggered Normalization

    Thu Oct 8Grand Ballroom #107All about training
  • Towards Full Pipeline FP8 Reinforcement Learning for LLMs

    Thu Oct 8Grand Ballroom #108Compute-efficient LMs
  • Safe Inference-Time Alignment via Lagrangian Reward Augmentation

    Thu Oct 8Grand Ballroom #109All about safety
  • RL-VLA^3: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

    Thu Oct 8Grand Ballroom #110LMs and embodiment
  • Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning

    Thu Oct 8Grand Ballroom #111All about training
  • ExpRL: Exploratory RL for LLM Mid-Training

    Thu Oct 8Grand Ballroom #112All about training
  • Diversity or Precision? A Deep Dive into Next Token Prediction

    ORALThu Oct 8Grand Ballroom #113Science of LMs
  • Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

    Thu Oct 8Grand Ballroom #114All about training
  • Token-Level Control or Just a Better Mean? Isolating the Gains of Adaptive Decoding

    Thu Oct 8Grand Ballroom #115Inference algorithms for LMs
  • Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

    Thu Oct 8Grand Ballroom #116LMs and interactions
  • Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics

    Thu Oct 8Grand Ballroom #117All about training
  • Understanding Machine Unlearning Through the Lens of Mode Connectivity

    Thu Oct 8Grand Ballroom #118Learning algorithms for LMs
  • On the Impossibility of Retrain Equivalence in Machine Unlearning

    Thu Oct 8Grand Ballroom #119Learning algorithms for LMs
  • LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

    Thu Oct 8Grand Ballroom #120Learning algorithms for LMs
  • MoRFI: Monotonic Sparse Autoencoder Feature Identification

    Thu Oct 8Grand Ballroom #122Science of LMs
  • UNMASK Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

    Thu Oct 8Grand Ballroom #123Science of LMs
  • Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts

    Thu Oct 8Grand Ballroom #124All about evaluation
  • TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models

    Thu Oct 8Grand Ballroom #125Diverse domains & novel applications
  • CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs

    Thu Oct 8Grand Ballroom #126Multimodal LMs
  • The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

    Thu Oct 8Grand Ballroom #127Multimodal LMs
  • Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

    Thu Oct 8Grand Ballroom #128Multimodal LMs
  • CoVLA: Vision-Language Alignment with Fewer Vision Tokens

    Thu Oct 8Grand Ballroom #129Multimodal LMs
  • EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

    Thu Oct 8Grand Ballroom #130All about data
  • Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

    Thu Oct 8Grand Ballroom #131All about training
  • EvoSkillBank: Hierarchical Skill Self-Evolution and Skill-Bank Governance for Continual Agent Learning

    Thu Oct 8Grand Ballroom #132Learning algorithms for LMs
  • Toward Scalable Terminal Task Synthesis via Skill Graphs

    Thu Oct 8Grand Ballroom #133LMs with tools and code
  • Budget-Aware Tool Use Enables Effective Agent Scaling

    Thu Oct 8Grand Ballroom #134LMs with tools and code
  • Thought-Level Beam Search for Reasoning

    Thu Oct 8Grand Ballroom #135Compute-efficient LMs
  • TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

    Thu Oct 8Grand Ballroom #136Compute-efficient LMs
  • HarmThoughts: A Benchmark for Fine-Grained Harmful Behavior Detection in Reasoning Traces

    Thu Oct 8Grand Ballroom #137All about safety
  • TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking

    Thu Oct 8Grand Ballroom #138All about safety
  • Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

    Thu Oct 8Grand Ballroom #139All about safety
  • PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

    Thu Oct 8Grand Ballroom #140All about safety
  • Improved Robustness against Indirect Prompt Injection Can Be Built into LLMs

    Thu Oct 8Grand Ballroom #141All about safety
  • Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense

    Thu Oct 8Grand Ballroom #142All about safety
  • MEMENTO: Teaching LLMs to Manage Their Context

    ORALThu Oct 8Grand Ballroom #143Compute-efficient LMs
  • What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

    ORALThu Oct 8Grand Ballroom #144All about evaluation
  • Introspective Diffusion Language Models

    ORALThu Oct 8Grand Ballroom #145Learning algorithms for LMs
<!-- ==== Scripts (order matters:shell → data → page)==== -->