← People
Antoine Bosselut
following
EPFL
@ABosselut
Papers in the feed →
Papers · 53
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
2026-07-05
alphaXiv
arXiv
S2
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
arXiv.org
2026-05-29
alphaXiv
arXiv
S2
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
arXiv.org
2026-04-21
alphaXiv
arXiv
S2
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
arXiv.org
2026-04-14
alphaXiv
arXiv
S2
AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
International Conference on Human Factors in Computing Systems
2026-04-13
S2
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
arXiv.org
2026-04-03
alphaXiv
arXiv
S2
Large Language Models Align with the Human Brain during Creative Thinking
arXiv.org
2026-04-03
alphaXiv
arXiv
S2
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
arXiv.org
2026-03-09
alphaXiv
arXiv
S2
Brittlebench: Quantifying LLM robustness via prompt sensitivity
arXiv.org
2026-02-27
alphaXiv
arXiv
S2
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
arXiv.org
2026-02-18
alphaXiv
arXiv
S2
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
Conference of the European Chapter of the Association for Computational Linguistics
2026-01-21
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Annual Meeting of the Association for Computational Linguistics
2026
S2
DRIVINGVQA: A Dataset for Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
Conference of the European Chapter of the Association for Computational Linguistics
2026
S2
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Advances in Neural Information Processing Systems 38
2025-11-03
alphaXiv
arXiv
S2
Revisiting Multilingual Data Mixtures in Language Model Pretraining
arXiv.org
2025-10-29
alphaXiv
arXiv
S2
RLMEval: Evaluating Research-Level Neural Theorem Proving
Conference on Empirical Methods in Natural Language Processing
2025-10-29
alphaXiv
arXiv
S2
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
Conference on Empirical Methods in Natural Language Processing
2025-10-29
alphaXiv
arXiv
S2
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
arXiv.org
2025-10-22
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
2025-09-17
alphaXiv
arXiv
S2
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
Annual Meeting of the Association for Computational Linguistics
2025-09-05
alphaXiv
arXiv
S2
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Annual Meeting of the Association for Computational Linguistics
2025-08-06
alphaXiv
arXiv
S2
GeoExplorer: Active Geo-Localization with Curiosity-Driven Exploration
IEEE International Conference on Computer Vision
2025-07-31
alphaXiv
arXiv
S2
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Conference on Empirical Methods in Natural Language Processing
2025-07-11
alphaXiv
arXiv
S2
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
arXiv.org
2025-07-08
alphaXiv
arXiv
S2
Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison
Workshop on Innovative Use of NLP for Building Educational Applications
2025-07-02
alphaXiv
arXiv
S2
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
Conference of the European Chapter of the Association for Computational Linguistics
2025-06-18
alphaXiv
arXiv
S2
Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
arXiv.org
2025-06-16
alphaXiv
arXiv
S2
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
arXiv.org
2025-06-09
alphaXiv
arXiv
S2
Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings
Conference on Empirical Methods in Natural Language Processing
2025-05-30
alphaXiv
arXiv
S2
Creative Preference Optimization
Conference on Empirical Methods in Natural Language Processing
2025-05-20
alphaXiv
arXiv
S2
Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
Neural Information Processing Systems
2025-05-19
alphaXiv
arXiv
S2
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
arXiv.org
2025-04-08
alphaXiv
arXiv
S2
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
Computer Vision and Pattern Recognition
2025-03-26
alphaXiv
arXiv
S2
From Language to Cognition: How LLMs Outgrow the Human Language Network
Conference on Empirical Methods in Natural Language Processing
2025-03-03
alphaXiv
arXiv
S2
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
2025-01-08
alphaXiv
arXiv
S2
For Better or for Worse, Transformers Seek Patterns for Memorization
Neural Information Processing Systems
2025
S2
DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests
arXiv.org
2025
S2
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Annual Meeting of the Association for Computational Linguistics
2025
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
arXiv.org
2025
S2
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
North American Chapter of the Association for Computational Linguistics
2024-12-16
alphaXiv
arXiv
S2
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
arXiv.org
2024-12-04
alphaXiv
arXiv
S2
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
International Conference on Learning Representations
2024-11-29
alphaXiv
arXiv
S2
The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
North American Chapter of the Association for Computational Linguistics
2024-11-04
alphaXiv
arXiv
S2
Creativity in AI: Progresses and Challenges
arXiv.org
2024-10-22
alphaXiv
arXiv
S2
Evaluating Morphological Compositional Generalization in Large Language Models
North American Chapter of the Association for Computational Linguistics
2024-10-16
alphaXiv
arXiv
S2
LLMs Are In-Context Bandit Reinforcement Learners
2024-10-07
alphaXiv
arXiv
S2
EPFL-MAKE at “Discharge Me!”: An LLM System for Automatically Generating Discharge Summaries of Clinical Electronic Health Record
Workshop on Biomedical Natural Language Processing
2024
S2
Realized Risk Reduction in Portfolio Optimization through Iterative Solutions of the Mvdr Filter
S2
Fine-tuning Vision-Language Models for Animal Behavior Analysis
S2
Cross-Lingual Multi-Hop Knowledge Editing – Benchmarks, Analysis and a Simple Contrastive Learning based Approach
S2
Graph Memory-based Editing for Large Language Models
S2
Introductory Tutorial: Commonsense Reasoning for Natural Language Processing
S2
Upper-bound Translation Performance of Llama-2 Under Idealized Setup
S2