← People
Sara Hooker

Sara Hooker following

Cohere Labs
@sarahookrPapers in the feed →

Papers · 42
  1. Open-World Evaluations for Measuring Frontier AI Capabilities
    arXiv.org2026-05-19alphaXiv arXiv S2
  2. Tiny Aya: Bridging Scale and Multilingual Depth
    arXiv.org2026-03-12alphaXiv arXiv S2
  3. SimMerge: Learning to Select Merge Operators from Similarity Signals
    arXiv.org2026-01-14alphaXiv arXiv S2
  4. The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
    arXiv.org2025-10-22alphaXiv arXiv S2
  5. The Disparate Impacts of Speculative Decoding
    arXiv.org2025-10-02alphaXiv arXiv S2
  6. Verification Limits Code LLM Training
    arXiv.org2025-09-25alphaXiv arXiv S2
  7. When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
    Conference on Empirical Methods in Natural Language Processing2025-06-25alphaXiv arXiv S2
  8. Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
    arXiv.org2025-06-18alphaXiv arXiv S2
  9. Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers
    Neural Information Processing Systems2025-06-17alphaXiv arXiv S2
  10. One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
    Annual Meeting of the Association for Computational Linguistics2025-06-12alphaXiv arXiv S2
  11. The Multilingual Divide and Its Impact on Global AI Safety
    arXiv.org2025-05-27alphaXiv arXiv S2
  12. Aya Vision: Advancing the Frontier of Multilingual Multimodality
    arXiv.org2025-05-13alphaXiv arXiv S2
  13. The Leaderboard Illusion
    Neural Information Processing Systems2025-04-29alphaXiv arXiv S2
  14. Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
    arXiv.org2025-04-09alphaXiv arXiv S2
  15. Command A: An Enterprise-Ready Large Language Model
    2025-04-01alphaXiv arXiv S2
  16. MMTEB: Massive Multilingual Text Embedding Benchmark
    arXiv.org2025-02-19alphaXiv arXiv S2
  17. Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
    North American Chapter of the Association for Computational Linguistics2025-02-04alphaXiv arXiv S2
  18. Fairness of Deep Ensembles: On the interplay between per-group task difficulty and under-representation
    Conference on Fairness, Accountability and Transparency2025-01-24alphaXiv arXiv S2
  19. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    Annual Meeting of the Association for Computational Linguistics2025S2
  20. RLHF Algorithms Ranked: An Extensive Evaluation Across Diverse Tasks, Rewards, and Hyperparameters
    Conference on Empirical Methods in Natural Language Processing2025S2
  21. To Code or Not To Code? Exploring Impact of Code in Pre-training
    International Conference on Learning Representations2025S2
  22. Nexus: Adaptive Upcycling to Efficiently Pretrain Mixture of Experts
    Conference on Empirical Methods in Natural Language Processing2025S2
  23. Multilingual Arbitration: Optimizing Data Pools to Accelerate Multilingual Progress
    Annual Meeting of the Association for Computational Linguistics2025S2
  24. Open Problems in Technical AI Governance
    Trans. Mach. Learn. Res.2025S2
  25. Bridging the Data Provenance Gap Across Text, Speech and Video
    arXiv.org2024-12-19alphaXiv arXiv S2
  26. Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
    arXiv.org2024-12-05alphaXiv arXiv S2
  27. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    arXiv.org2024-12-04alphaXiv arXiv S2
  28. The Reality of AI and Biorisk
    Conference on Fairness, Accountability and Transparency2024-12-02alphaXiv arXiv S2
  29. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
    International Conference on Learning Representations2024-11-29alphaXiv arXiv S2
  30. M-RewardBench: Evaluating Reward Models in Multilingual Settings
    Annual Meeting of the Association for Computational Linguistics2024-10-20alphaXiv arXiv S2
  31. Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
    arXiv.org2024-10-14alphaXiv arXiv S2
  32. Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
    arXiv.org2024-08-28alphaXiv arXiv S2
  33. Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
    arXiv.org2024-08-27alphaXiv arXiv S2
  34. To Code, or Not To Code? Exploring Impact of Code in Pre-training
    arXiv.org2024-08-20alphaXiv arXiv S2
  35. The future of open human feedback
    Nature Machine Intelligence2024-08-15alphaXiv arXiv S2
  36. LLM See, LLM Do: Leveraging Active Inheritance to Target Non-Differentiable Objectives
    Conference on Empirical Methods in Natural Language Processing2024S2
  37. Robust distillation for worst-class performance: on the interplay between teacher and student objectives
    Conference on Uncertainty in Artificial Intelligence2023S2
  38. The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
    Neural Information Processing Systems2023S2
  39. Cross-lingual Transfer Dynamics in BLOOMZ: Insights into Multilingual Generalization
    S2
  40. Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
    S2
  41. The Data Provenance Project
    S2
  42. Capabilities and risks from frontier AI
    S2