← People
Antoine Bosselut

Antoine Bosselut following

EPFL
@ABosselutPapers in the feed →

Papers · 53
  1. DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
    2026-07-05alphaXiv arXiv S2
  2. Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
    arXiv.org2026-05-29alphaXiv arXiv S2
  3. Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
    arXiv.org2026-04-21alphaXiv arXiv S2
  4. An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
    arXiv.org2026-04-14alphaXiv arXiv S2
  5. AI meets Mathematics Education: Supporting Instructors in Large Mathematics Classes with Context-Aware AI
    International Conference on Human Factors in Computing Systems2026-04-13S2
  6. CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
    arXiv.org2026-04-03alphaXiv arXiv S2
  7. Large Language Models Align with the Human Brain during Creative Thinking
    arXiv.org2026-04-03alphaXiv arXiv S2
  8. AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
    arXiv.org2026-03-09alphaXiv arXiv S2
  9. Brittlebench: Quantifying LLM robustness via prompt sensitivity
    arXiv.org2026-02-27alphaXiv arXiv S2
  10. Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents
    arXiv.org2026-02-18alphaXiv arXiv S2
  11. Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
    Conference of the European Chapter of the Association for Computational Linguistics2026-01-21alphaXiv arXiv S2
  12. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
    Annual Meeting of the Association for Computational Linguistics2026S2
  13. DRIVINGVQA: A Dataset for Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
    Conference of the European Chapter of the Association for Computational Linguistics2026S2
  14. Measuring what Matters: Construct Validity in Large Language Model Benchmarks
    Advances in Neural Information Processing Systems 382025-11-03alphaXiv arXiv S2
  15. Revisiting Multilingual Data Mixtures in Language Model Pretraining
    arXiv.org2025-10-29alphaXiv arXiv S2
  16. RLMEval: Evaluating Research-Level Neural Theorem Proving
    Conference on Empirical Methods in Natural Language Processing2025-10-29alphaXiv arXiv S2
  17. CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
    Conference on Empirical Methods in Natural Language Processing2025-10-29alphaXiv arXiv S2
  18. GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
    arXiv.org2025-10-22alphaXiv arXiv S2
  19. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
    2025-09-17alphaXiv arXiv S2
  20. Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
    Annual Meeting of the Association for Computational Linguistics2025-09-05alphaXiv arXiv S2
  21. Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
    Annual Meeting of the Association for Computational Linguistics2025-08-06alphaXiv arXiv S2
  22. GeoExplorer: Active Geo-Localization with Curiosity-Driven Exploration
    IEEE International Conference on Computer Vision2025-07-31alphaXiv arXiv S2
  23. From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation
    Conference on Empirical Methods in Natural Language Processing2025-07-11alphaXiv arXiv S2
  24. PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
    arXiv.org2025-07-08alphaXiv arXiv S2
  25. Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison
    Workshop on Innovative Use of NLP for Building Educational Applications2025-07-02alphaXiv arXiv S2
  26. ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
    Conference of the European Chapter of the Association for Computational Linguistics2025-06-18alphaXiv arXiv S2
  27. Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
    arXiv.org2025-06-16alphaXiv arXiv S2
  28. AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
    arXiv.org2025-06-09alphaXiv arXiv S2
  29. Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings
    Conference on Empirical Methods in Natural Language Processing2025-05-30alphaXiv arXiv S2
  30. Creative Preference Optimization
    Conference on Empirical Methods in Natural Language Processing2025-05-20alphaXiv arXiv S2
  31. Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks
    Neural Information Processing Systems2025-05-19alphaXiv arXiv S2
  32. Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
    arXiv.org2025-04-08alphaXiv arXiv S2
  33. VinaBench: Benchmark for Faithful and Consistent Visual Narratives
    Computer Vision and Pattern Recognition2025-03-26alphaXiv arXiv S2
  34. From Language to Cognition: How LLMs Outgrow the Human Language Network
    Conference on Empirical Methods in Natural Language Processing2025-03-03alphaXiv arXiv S2
  35. Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
    2025-01-08alphaXiv arXiv S2
  36. For Better or for Worse, Transformers Seek Patterns for Memorization
    Neural Information Processing Systems2025S2
  37. DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests
    arXiv.org2025S2
  38. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    Annual Meeting of the Association for Computational Linguistics2025S2
  39. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
    arXiv.org2025S2
  40. PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
    North American Chapter of the Association for Computational Linguistics2024-12-16alphaXiv arXiv S2
  41. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    arXiv.org2024-12-04alphaXiv arXiv S2
  42. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
    International Conference on Learning Representations2024-11-29alphaXiv arXiv S2
  43. The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units
    North American Chapter of the Association for Computational Linguistics2024-11-04alphaXiv arXiv S2
  44. Creativity in AI: Progresses and Challenges
    arXiv.org2024-10-22alphaXiv arXiv S2
  45. Evaluating Morphological Compositional Generalization in Large Language Models
    North American Chapter of the Association for Computational Linguistics2024-10-16alphaXiv arXiv S2
  46. LLMs Are In-Context Bandit Reinforcement Learners
    2024-10-07alphaXiv arXiv S2
  47. EPFL-MAKE at “Discharge Me!”: An LLM System for Automatically Generating Discharge Summaries of Clinical Electronic Health Record
    Workshop on Biomedical Natural Language Processing2024S2
  48. Realized Risk Reduction in Portfolio Optimization through Iterative Solutions of the Mvdr Filter
    S2
  49. Fine-tuning Vision-Language Models for Animal Behavior Analysis
    S2
  50. Cross-Lingual Multi-Hop Knowledge Editing – Benchmarks, Analysis and a Simple Contrastive Learning based Approach
    S2
  51. Graph Memory-based Editing for Large Language Models
    S2
  52. Introductory Tutorial: Commonsense Reasoning for Natural Language Processing
    S2
  53. Upper-bound Translation Performance of Llama-2 Under Idealized Setup
    S2