← People
Tim Rocktäschel
following
UCL / Google DeepMind
@_rockt
Papers in the feed →
Papers · 24
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
arXiv.org
2026-06-06
alphaXiv
arXiv
S2
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
arXiv.org
2026-01-05
alphaXiv
arXiv
S2
Check Your Work: Structured Checklist Feedback for Improving Large Language Models
Annual Meeting of the Association for Computational Linguistics
2026
S2
ACRM: Multi-Agent Trajectory Learning for Automated Credit Risk Model Refreshing in Production
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
2026
S2
Towards Uncovering How Large Language ModelsWork: An Interpretability Perspective
SIGKDD Explorations
2025-12-30
S2
Imagined Autocurricula
Neural Information Processing Systems
2025-09-11
alphaXiv
arXiv
S2
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
arXiv.org
2025-09-03
alphaXiv
arXiv
S2
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
2025-06-23
alphaXiv
arXiv
S2
LLM-First Search: Self-Guided Exploration of the Solution Space
arXiv.org
2025-06-05
alphaXiv
arXiv
S2
D3PO: Preference-Based Alignment of Discrete Diffusion Models
arXiv.org
2025-03-11
alphaXiv
arXiv
S2
Investigating Non-Transitivity in LLM-as-a-Judge
International Conference on Machine Learning
2025-02-19
alphaXiv
arXiv
S2
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models with Human Feedback
Conference on Empirical Methods in Natural Language Processing
2025
S2
Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training
arXiv.org
2025
S2
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
International Conference on Learning Representations
2025
S2
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
International Conference on Learning Representations
2024-11-20
alphaXiv
arXiv
S2
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
arXiv.org
2024-11-19
alphaXiv
arXiv
S2
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
arXiv.org
2024-10-04
alphaXiv
arXiv
S2
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
International Conference on Learning Representations
2024
S2
The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
Neural Information Processing Systems
2023
S2
JaxMARL: Multi-Agent RL Environments in JAX
arXiv.org
2023
S2
Graph Memory-based Editing for Large Language Models
S2
Conference on Empirical Methods in Natural Language Processing Proceedings of the Fourth International Workshop on Natural Language Processing for Social Media Socialnlp@emnlp2016 Chairs' Welcome Identifying and Categorizing Disaster-related Tweets Why Do They Leave: Modeling Participation in Online
S2
A Systematic Literature Review of Adapter-based Approaches to Knowledge-enhanced Language Models
S2
On Reward Functions For Self-Improving Chain-of-Thought Reasoning Without Supervised Datasets (Abridged Version)
S2