← People
Graham Neubig

Graham Neubig following

Carnegie Mellon University
@gneubigPapers in the feed →

Papers · 124
  1. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
    Annual Meeting of the Association for Computational Linguistics2026-07-08alphaXiv arXiv S2
  2. PACE: A Proxy for Agentic Capability Evaluation
    2026-07-02alphaXiv arXiv S2
  3. Decomposer: Learning to Decompile Symbolic Music to Programs
    2026-07-02alphaXiv arXiv S2
  4. PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
    arXiv.org2026-06-30alphaXiv arXiv S2
  5. Discretizing Reward Models
    arXiv.org2026-06-19alphaXiv arXiv S2
  6. Enhancing Software Engineering Through Closed-Loop Memory Optimization
    arXiv.org2026-06-04alphaXiv arXiv S2
  7. On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
    arXiv.org2026-05-20alphaXiv arXiv S2
  8. Reinforcing Human Behavior Simulation via Verbal Feedback
    arXiv.org2026-05-19alphaXiv arXiv S2
  9. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
    arXiv.org2026-05-09alphaXiv arXiv S2
  10. Recursive Agent Optimization
    arXiv.org2026-05-07alphaXiv arXiv S2
  11. Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
    arXiv.org2026-04-16alphaXiv arXiv S2
  12. What do Language Models Learn and When? The Implicit Curriculum Hypothesis
    arXiv.org2026-04-09alphaXiv arXiv S2
  13. Gym-Anything: Turn any Software into an Agent Environment
    arXiv.org2026-04-07alphaXiv arXiv S2
  14. IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
    arXiv.org2026-04-06alphaXiv arXiv S2
  15. Effective Strategies for Asynchronous Software Engineering Agents
    arXiv.org2026-03-23alphaXiv arXiv S2
  16. Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
    arXiv.org2026-03-19alphaXiv arXiv S2
  17. CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
    arXiv.org2026-03-18alphaXiv arXiv S2
  18. CUBE: A Standard for Unifying Agent Benchmarks
    arXiv.org2026-03-16alphaXiv arXiv S2
  19. Mind the Sim2Real Gap in User Simulation for Agentic Tasks
    arXiv.org2026-03-11alphaXiv arXiv S2
  20. A Rubric-Supervised Critic from Sparse Real-World Outcomes
    arXiv.org2026-03-04alphaXiv arXiv S2
  21. Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
    arXiv.org2026-03-03alphaXiv arXiv S2
  22. How Well Does Agent Development Reflect Real-World Work?
    arXiv.org2026-03-01alphaXiv arXiv S2
  23. Modeling Distinct Human Interaction in Web Agents
    arXiv.org2026-02-19alphaXiv arXiv S2
  24. Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
    arXiv.org2026-02-18alphaXiv arXiv S2
  25. Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
    Annual Meeting of the Association for Computational Linguistics2026-01-26alphaXiv arXiv S2
  26. Massively Multilingual Joint Segmentation and Glossing
    Annual Meeting of the Association for Computational Linguistics2026-01-16alphaXiv arXiv S2
  27. CAIRE: Cultural Attribution of Images with Retrieval
    Conference of the European Chapter of the Association for Computational Linguistics2026S2
  28. Training Versatile Coding Agents in Synthetic Environments
    arXiv.org2025-12-13alphaXiv arXiv S2
  29. On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
    arXiv.org2025-12-08alphaXiv arXiv S2
  30. ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
    arXiv.org2025-12-04alphaXiv arXiv S2
  31. RefineBench: Evaluating Refinement Capability of Language Models via Checklists
    arXiv.org2025-11-27alphaXiv arXiv S2
  32. Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
    IEEE Workshop/Winter Conference on Applications of Computer Vision2025-11-18alphaXiv arXiv S2
  33. EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
    arXiv.org2025-11-06alphaXiv arXiv S2
  34. The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
    arXiv.org2025-11-05alphaXiv arXiv S2
  35. Training Proactive and Personalized LLM Agents
    arXiv.org2025-11-04alphaXiv arXiv S2
  36. Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
    arXiv.org2025-11-04alphaXiv arXiv S2
  37. Accumulating Context Changes the Beliefs of Language Models
    arXiv.org2025-11-03alphaXiv arXiv S2
  38. The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
    arXiv.org2025-10-29alphaXiv arXiv S2
  39. Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
    arXiv.org2025-10-28alphaXiv arXiv S2
  40. How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
    arXiv.org2025-10-26alphaXiv arXiv S2
  41. TOM-SWE: User Mental Modeling For Software Engineering Agents
    arXiv.org2025-10-24alphaXiv arXiv S2
  42. TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
    arXiv.org2025-10-22alphaXiv arXiv S2
  43. Prompt-MII: Meta-Learning Instruction Induction for LLMs
    arXiv.org2025-10-19alphaXiv arXiv S2
  44. Midtraining Bridges Pretraining and Posttraining Distributions
    arXiv.org2025-10-16alphaXiv arXiv S2
  45. MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
    Transactions of the Association for Computational Linguistics2025-10-16alphaXiv arXiv S2
  46. How can we assess human-agent interactions? Case studies in software agent design
    arXiv.org2025-10-10alphaXiv arXiv S2
  47. Grounding Multilingual Multimodal LLMs With Cultural Knowledge
    Conference on Empirical Methods in Natural Language Processing2025-08-10alphaXiv arXiv S2
  48. Devstral: Fine-tuning Language Models for Coding Agent Applications
    arXiv.org2025-08-08alphaXiv arXiv S2
  49. Checklists Are Better Than Reward Models For Aligning Language Models
    Neural Information Processing Systems2025-07-24alphaXiv arXiv S2
  50. Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
    International Conference on Human Factors in Computing Systems2025-07-10alphaXiv arXiv S2
  51. OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
    arXiv.org2025-07-08alphaXiv arXiv S2
  52. Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
    arXiv.org2025-07-01alphaXiv arXiv S2
  53. ZINA: Multimodal Fine-grained Hallucination Detection and Editing
    arXiv.org2025-06-16alphaXiv arXiv S2
  54. CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
    arXiv.org2025-06-10alphaXiv arXiv S2
  55. Go-Browse: Training Web Agents with Structured Exploration
    arXiv.org2025-06-04alphaXiv arXiv S2
  56. Coding Agents with Multimodal Browsing are Generalist Problem Solvers
    Conference of the European Chapter of the Association for Computational Linguistics2025-06-03alphaXiv arXiv S2
  57. BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
    Annual Meeting of the Association for Computational Linguistics2025-06-02alphaXiv arXiv S2
  58. FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
    arXiv.org2025-05-26alphaXiv arXiv S2
  59. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
    arXiv.org2025-05-15alphaXiv arXiv S2
  60. VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
    arXiv.org2025-04-14alphaXiv arXiv S2
  61. Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
    arXiv.org2025-04-10alphaXiv arXiv S2
  62. SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
    arXiv.org2025-04-09alphaXiv arXiv S2
  63. Inducing Programmatic Skills for Agentic Tasks
    arXiv.org2025-04-09alphaXiv arXiv S2
  64. M-Prometheus: A Suite of Open Multilingual LLM Judges
    arXiv.org2025-04-07alphaXiv arXiv S2
  65. Scaling Evaluation-Time Compute with Reasoning Models as Evaluators
    Annual Meeting of the Association for Computational Linguistics2025-03-25alphaXiv arXiv S2
  66. Overtrained Language Models Are Harder to Fine-Tune
    International Conference on Machine Learning2025-03-24alphaXiv arXiv S2
  67. Benchmarking Failures in Tool-Augmented Language Models
    North American Chapter of the Association for Computational Linguistics2025-03-18alphaXiv arXiv S2
  68. Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
    Annual Meeting of the Association for Computational Linguistics2025-03-11alphaXiv arXiv S2
  69. Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
    Conference on Empirical Methods in Natural Language Processing2025-03-05alphaXiv arXiv S2
  70. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
    North American Chapter of the Association for Computational Linguistics2025-02-21alphaXiv arXiv S2
  71. Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
    2025-02-18alphaXiv arXiv S2
  72. The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
    arXiv.org2025-02-12alphaXiv arXiv S2
  73. Demystifying Long Chain-of-Thought Reasoning in LLMs
    arXiv.org2025-02-05alphaXiv arXiv S2
  74. Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
    North American Chapter of the Association for Computational Linguistics2025-02-04alphaXiv arXiv S2
  75. CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
    North American Chapter of the Association for Computational Linguistics2025-01-28alphaXiv arXiv S2
  76. AutoPresent: Designing Structured Visuals from Scratch
    Computer Vision and Pattern Recognition2025-01-01alphaXiv arXiv S2
  77. Evaluating Numeracy of Language Models as a Natural Language Inference Task
    North American Chapter of the Association for Computational Linguistics2025S2
  78. AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories
    Conference on Empirical Methods in Natural Language Processing2025S2
  79. Demystifying Long Chain-of-Thought Reasoning
    International Conference on Machine Learning2025S2
  80. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
    Annual Meeting of the Association for Computational Linguistics2025S2
  81. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
    International Conference on Learning Representations2025S2
  82. Synthetic Data in the Era of Large Language Models
    Annual Meeting of the Association for Computational Linguistics2025S2
  83. Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators
    arXiv.org2025S2
  84. Training Software Engineering Agents and Verifiers with SWE-Gym
    International Conference on Machine Learning2024-12-30alphaXiv arXiv S2
  85. TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
    Advances in Neural Information Processing Systems 382024-12-18alphaXiv arXiv S2
  86. Towards Automatic Evaluation for Image Transcreation
    North American Chapter of the Association for Computational Linguistics2024-12-18alphaXiv arXiv S2
  87. HILITE: Human-in-the-loop Interactive Tool for Image Editing
    BigData Congress [Services Society]2024-12-15S2
  88. MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
    arXiv.org2024-12-06alphaXiv arXiv S2
  89. The BrowserGym Ecosystem for Web Agent Research
    Trans. Mach. Learn. Res.2024-12-06alphaXiv arXiv S2
  90. Evaluating Language Models as Synthetic Data Generators
    Annual Meeting of the Association for Computational Linguistics2024-12-04alphaXiv arXiv S2
  91. OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
    arXiv.org2024-11-21alphaXiv arXiv S2
  92. What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
    North American Chapter of the Association for Computational Linguistics2024-11-04alphaXiv arXiv S2
  93. JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
    North American Chapter of the Association for Computational Linguistics2024-10-22alphaXiv arXiv S2
  94. Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
    arXiv.org2024-10-21alphaXiv arXiv S2
  95. Beyond Browsing: API-Based Web Agents
    Annual Meeting of the Association for Computational Linguistics2024-10-21alphaXiv arXiv S2
  96. NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
    Neural Information Processing Systems2024-10-18alphaXiv arXiv S2
  97. Harnessing Webpage UIs for Text-Rich Visual Understanding
    International Conference on Learning Representations2024-10-17alphaXiv arXiv S2
  98. Stereotype or Personalization? User Identity Biases Chatbot Recommendations
    Annual Meeting of the Association for Computational Linguistics2024-10-08alphaXiv arXiv S2
  99. Better Instruction-Following Through Minimum Bayes Risk
    International Conference on Learning Representations2024-10-03alphaXiv arXiv S2
  100. Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
    Neural Information Processing Systems2024-09-24alphaXiv arXiv S2
  101. Agent Workflow Memory
    International Conference on Machine Learning2024-09-11alphaXiv arXiv S2
  102. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
    Annual Meeting of the Association for Computational Linguistics2024-09-04alphaXiv arXiv S2
  103. Program-Aided Reasoners (Better) Know What They Know
    North American Chapter of the Association for Computational Linguistics2024S2
  104. VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
    Annual Meeting of the Association for Computational Linguistics2024S2
  105. An image speaks a thousand words, but can everyone listen? On translating images for cultural relevance
    arXiv.org2024S2
  106. Alignment for Honesty
    Advances in Neural Information Processing Systems 372024S2
  107. An Incomplete Loop: Deductive, Inductive, and Abductive Learning in Large Language Models
    arXiv.org2024S2
  108. OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
    arXiv.org2024S2
  109. RAGGED: Towards Informed Design of Retrieval Augmented Generation Systems
    arXiv.org2024S2
  110. CMU’s IWSLT 2024 Offline Speech Translation System: A Cascaded Approach For Long-Form Robustness
    International Workshop on Spoken Language Translation2024S2
  111. Divergences between Language Models and Human Brains
    Neural Information Processing Systems2024S2
  112. GlossLM: Multilingual Pretraining for Low-Resource Interlinear Glossing
    arXiv.org2024S2
  113. DiffusER: Diffusion via Edit-based Reconstruction
    International Conference on Learning Representations2023S2
  114. SigMoreFun Submission to the SIGMORPHON Shared Task on Interlinear Glossing
    Special Interest Group on Computational Morphology and Phonology Workshop2023S2
  115. Recent Developments in Computational Typology and Multilingual Natural Language Processing
    S2
  116. Towards Scalable Oversight: Meta-Evaluation of LLMs as Evaluators via Agent Debate
    S2
  117. Dynamic Summarization Length Control via Model Merging
    S2
  118. A Survey of Cross-Lingual Alignment: Definitions, Methods, Future
    S2
  119. CoderGen: Towards Domain-Specific Code Generation of Large Language Models
    S2
  120. GenAI-Bench: A Holistic Benchmark for Compositional Text-to-Visual Generation
    S2
  121. P OSITION : A GENTIC S YSTEMS S HOULD BE G ENERAL
    S2
  122. Position: Humans are Missing from AI Coding Agent Research
    S2
  123. Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
    S2
  124. Multi-Task Learning with Self-Supervised Objectives can Improve Worst-Group Outcomes
    S2