← People
Graham Neubig
following
Carnegie Mellon University
@gneubig
Papers in the feed →
Papers · 124
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Annual Meeting of the Association for Computational Linguistics
2026-07-08
alphaXiv
arXiv
S2
PACE: A Proxy for Agentic Capability Evaluation
2026-07-02
alphaXiv
arXiv
S2
Decomposer: Learning to Decompile Symbolic Music to Programs
2026-07-02
alphaXiv
arXiv
S2
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
arXiv.org
2026-06-30
alphaXiv
arXiv
S2
Discretizing Reward Models
arXiv.org
2026-06-19
alphaXiv
arXiv
S2
Enhancing Software Engineering Through Closed-Loop Memory Optimization
arXiv.org
2026-06-04
alphaXiv
arXiv
S2
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
arXiv.org
2026-05-20
alphaXiv
arXiv
S2
Reinforcing Human Behavior Simulation via Verbal Feedback
arXiv.org
2026-05-19
alphaXiv
arXiv
S2
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
arXiv.org
2026-05-09
alphaXiv
arXiv
S2
Recursive Agent Optimization
arXiv.org
2026-05-07
alphaXiv
arXiv
S2
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
arXiv.org
2026-04-16
alphaXiv
arXiv
S2
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
arXiv.org
2026-04-09
alphaXiv
arXiv
S2
Gym-Anything: Turn any Software into an Agent Environment
arXiv.org
2026-04-07
alphaXiv
arXiv
S2
IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
arXiv.org
2026-04-06
alphaXiv
arXiv
S2
Effective Strategies for Asynchronous Software Engineering Agents
arXiv.org
2026-03-23
alphaXiv
arXiv
S2
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
arXiv.org
2026-03-19
alphaXiv
arXiv
S2
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
arXiv.org
2026-03-18
alphaXiv
arXiv
S2
CUBE: A Standard for Unifying Agent Benchmarks
arXiv.org
2026-03-16
alphaXiv
arXiv
S2
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
arXiv.org
2026-03-11
alphaXiv
arXiv
S2
A Rubric-Supervised Critic from Sparse Real-World Outcomes
arXiv.org
2026-03-04
alphaXiv
arXiv
S2
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
arXiv.org
2026-03-03
alphaXiv
arXiv
S2
How Well Does Agent Development Reflect Real-World Work?
arXiv.org
2026-03-01
alphaXiv
arXiv
S2
Modeling Distinct Human Interaction in Web Agents
arXiv.org
2026-02-19
alphaXiv
arXiv
S2
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
arXiv.org
2026-02-18
alphaXiv
arXiv
S2
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
Annual Meeting of the Association for Computational Linguistics
2026-01-26
alphaXiv
arXiv
S2
Massively Multilingual Joint Segmentation and Glossing
Annual Meeting of the Association for Computational Linguistics
2026-01-16
alphaXiv
arXiv
S2
CAIRE: Cultural Attribution of Images with Retrieval
Conference of the European Chapter of the Association for Computational Linguistics
2026
S2
Training Versatile Coding Agents in Synthetic Environments
arXiv.org
2025-12-13
alphaXiv
arXiv
S2
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
arXiv.org
2025-12-08
alphaXiv
arXiv
S2
ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
arXiv.org
2025-12-04
alphaXiv
arXiv
S2
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
arXiv.org
2025-11-27
alphaXiv
arXiv
S2
Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
IEEE Workshop/Winter Conference on Applications of Computer Vision
2025-11-18
alphaXiv
arXiv
S2
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
arXiv.org
2025-11-06
alphaXiv
arXiv
S2
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
arXiv.org
2025-11-05
alphaXiv
arXiv
S2
Training Proactive and Personalized LLM Agents
arXiv.org
2025-11-04
alphaXiv
arXiv
S2
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
arXiv.org
2025-11-04
alphaXiv
arXiv
S2
Accumulating Context Changes the Beliefs of Language Models
arXiv.org
2025-11-03
alphaXiv
arXiv
S2
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
arXiv.org
2025-10-29
alphaXiv
arXiv
S2
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
arXiv.org
2025-10-28
alphaXiv
arXiv
S2
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
arXiv.org
2025-10-26
alphaXiv
arXiv
S2
TOM-SWE: User Mental Modeling For Software Engineering Agents
arXiv.org
2025-10-24
alphaXiv
arXiv
S2
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
arXiv.org
2025-10-22
alphaXiv
arXiv
S2
Prompt-MII: Meta-Learning Instruction Induction for LLMs
arXiv.org
2025-10-19
alphaXiv
arXiv
S2
Midtraining Bridges Pretraining and Posttraining Distributions
arXiv.org
2025-10-16
alphaXiv
arXiv
S2
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
Transactions of the Association for Computational Linguistics
2025-10-16
alphaXiv
arXiv
S2
How can we assess human-agent interactions? Case studies in software agent design
arXiv.org
2025-10-10
alphaXiv
arXiv
S2
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
Conference on Empirical Methods in Natural Language Processing
2025-08-10
alphaXiv
arXiv
S2
Devstral: Fine-tuning Language Models for Coding Agent Applications
arXiv.org
2025-08-08
alphaXiv
arXiv
S2
Checklists Are Better Than Reward Models For Aligning Language Models
Neural Information Processing Systems
2025-07-24
alphaXiv
arXiv
S2
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
International Conference on Human Factors in Computing Systems
2025-07-10
alphaXiv
arXiv
S2
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
arXiv.org
2025-07-08
alphaXiv
arXiv
S2
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
arXiv.org
2025-07-01
alphaXiv
arXiv
S2
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
arXiv.org
2025-06-16
alphaXiv
arXiv
S2
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
arXiv.org
2025-06-10
alphaXiv
arXiv
S2
Go-Browse: Training Web Agents with Structured Exploration
arXiv.org
2025-06-04
alphaXiv
arXiv
S2
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
Conference of the European Chapter of the Association for Computational Linguistics
2025-06-03
alphaXiv
arXiv
S2
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
Annual Meeting of the Association for Computational Linguistics
2025-06-02
alphaXiv
arXiv
S2
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
arXiv.org
2025-05-26
alphaXiv
arXiv
S2
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
arXiv.org
2025-05-15
alphaXiv
arXiv
S2
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
arXiv.org
2025-04-14
alphaXiv
arXiv
S2
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
arXiv.org
2025-04-10
alphaXiv
arXiv
S2
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
arXiv.org
2025-04-09
alphaXiv
arXiv
S2
Inducing Programmatic Skills for Agentic Tasks
arXiv.org
2025-04-09
alphaXiv
arXiv
S2
M-Prometheus: A Suite of Open Multilingual LLM Judges
arXiv.org
2025-04-07
alphaXiv
arXiv
S2
Scaling Evaluation-Time Compute with Reasoning Models as Evaluators
Annual Meeting of the Association for Computational Linguistics
2025-03-25
alphaXiv
arXiv
S2
Overtrained Language Models Are Harder to Fine-Tune
International Conference on Machine Learning
2025-03-24
alphaXiv
arXiv
S2
Benchmarking Failures in Tool-Augmented Language Models
North American Chapter of the Association for Computational Linguistics
2025-03-18
alphaXiv
arXiv
S2
Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
Annual Meeting of the Association for Computational Linguistics
2025-03-11
alphaXiv
arXiv
S2
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
Conference on Empirical Methods in Natural Language Processing
2025-03-05
alphaXiv
arXiv
S2
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
North American Chapter of the Association for Computational Linguistics
2025-02-21
alphaXiv
arXiv
S2
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
2025-02-18
alphaXiv
arXiv
S2
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
arXiv.org
2025-02-12
alphaXiv
arXiv
S2
Demystifying Long Chain-of-Thought Reasoning in LLMs
arXiv.org
2025-02-05
alphaXiv
arXiv
S2
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
North American Chapter of the Association for Computational Linguistics
2025-02-04
alphaXiv
arXiv
S2
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
North American Chapter of the Association for Computational Linguistics
2025-01-28
alphaXiv
arXiv
S2
AutoPresent: Designing Structured Visuals from Scratch
Computer Vision and Pattern Recognition
2025-01-01
alphaXiv
arXiv
S2
Evaluating Numeracy of Language Models as a Natural Language Inference Task
North American Chapter of the Association for Computational Linguistics
2025
S2
AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories
Conference on Empirical Methods in Natural Language Processing
2025
S2
Demystifying Long Chain-of-Thought Reasoning
International Conference on Machine Learning
2025
S2
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Annual Meeting of the Association for Computational Linguistics
2025
S2
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
International Conference on Learning Representations
2025
S2
Synthetic Data in the Era of Large Language Models
Annual Meeting of the Association for Computational Linguistics
2025
S2
Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators
arXiv.org
2025
S2
Training Software Engineering Agents and Verifiers with SWE-Gym
International Conference on Machine Learning
2024-12-30
alphaXiv
arXiv
S2
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Advances in Neural Information Processing Systems 38
2024-12-18
alphaXiv
arXiv
S2
Towards Automatic Evaluation for Image Transcreation
North American Chapter of the Association for Computational Linguistics
2024-12-18
alphaXiv
arXiv
S2
HILITE: Human-in-the-loop Interactive Tool for Image Editing
BigData Congress [Services Society]
2024-12-15
S2
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
arXiv.org
2024-12-06
alphaXiv
arXiv
S2
The BrowserGym Ecosystem for Web Agent Research
Trans. Mach. Learn. Res.
2024-12-06
alphaXiv
arXiv
S2
Evaluating Language Models as Synthetic Data Generators
Annual Meeting of the Association for Computational Linguistics
2024-12-04
alphaXiv
arXiv
S2
OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
arXiv.org
2024-11-21
alphaXiv
arXiv
S2
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
North American Chapter of the Association for Computational Linguistics
2024-11-04
alphaXiv
arXiv
S2
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
North American Chapter of the Association for Computational Linguistics
2024-10-22
alphaXiv
arXiv
S2
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
arXiv.org
2024-10-21
alphaXiv
arXiv
S2
Beyond Browsing: API-Based Web Agents
Annual Meeting of the Association for Computational Linguistics
2024-10-21
alphaXiv
arXiv
S2
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Neural Information Processing Systems
2024-10-18
alphaXiv
arXiv
S2
Harnessing Webpage UIs for Text-Rich Visual Understanding
International Conference on Learning Representations
2024-10-17
alphaXiv
arXiv
S2
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
Annual Meeting of the Association for Computational Linguistics
2024-10-08
alphaXiv
arXiv
S2
Better Instruction-Following Through Minimum Bayes Risk
International Conference on Learning Representations
2024-10-03
alphaXiv
arXiv
S2
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
Neural Information Processing Systems
2024-09-24
alphaXiv
arXiv
S2
Agent Workflow Memory
International Conference on Machine Learning
2024-09-11
alphaXiv
arXiv
S2
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Annual Meeting of the Association for Computational Linguistics
2024-09-04
alphaXiv
arXiv
S2
Program-Aided Reasoners (Better) Know What They Know
North American Chapter of the Association for Computational Linguistics
2024
S2
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
Annual Meeting of the Association for Computational Linguistics
2024
S2
An image speaks a thousand words, but can everyone listen? On translating images for cultural relevance
arXiv.org
2024
S2
Alignment for Honesty
Advances in Neural Information Processing Systems 37
2024
S2
An Incomplete Loop: Deductive, Inductive, and Abductive Learning in Large Language Models
arXiv.org
2024
S2
OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
arXiv.org
2024
S2
RAGGED: Towards Informed Design of Retrieval Augmented Generation Systems
arXiv.org
2024
S2
CMU’s IWSLT 2024 Offline Speech Translation System: A Cascaded Approach For Long-Form Robustness
International Workshop on Spoken Language Translation
2024
S2
Divergences between Language Models and Human Brains
Neural Information Processing Systems
2024
S2
GlossLM: Multilingual Pretraining for Low-Resource Interlinear Glossing
arXiv.org
2024
S2
DiffusER: Diffusion via Edit-based Reconstruction
International Conference on Learning Representations
2023
S2
SigMoreFun Submission to the SIGMORPHON Shared Task on Interlinear Glossing
Special Interest Group on Computational Morphology and Phonology Workshop
2023
S2
Recent Developments in Computational Typology and Multilingual Natural Language Processing
S2
Towards Scalable Oversight: Meta-Evaluation of LLMs as Evaluators via Agent Debate
S2
Dynamic Summarization Length Control via Model Merging
S2
A Survey of Cross-Lingual Alignment: Definitions, Methods, Future
S2
CoderGen: Towards Domain-Specific Code Generation of Large Language Models
S2
GenAI-Bench: A Holistic Benchmark for Compositional Text-to-Visual Generation
S2
P OSITION : A GENTIC S YSTEMS S HOULD BE G ENERAL
S2
Position: Humans are Missing from AI Coding Agent Research
S2
Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
S2
Multi-Task Learning with Self-Supervised Objectives can Improve Worst-Group Outcomes
S2