Marzieh Fadaee following
Cohere Labs
Papers · 43
-
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
-
AI Exposure Scores: what they measure, what they miss, and what comes next
-
The Culture Funnel: You Can't Align What isn't in the Data
-
Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards
-
Tiny Aya: Bridging Scale and Multilingual Depth
-
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
-
Unlocking Reasoning Capability on Machine Translation in Large Language Models
-
SimMerge: Learning to Select Merge Operators from Similarity Signals
-
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
-
Making, not Taking, the Best of N
-
Verification Limits Code LLM Training
-
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Conference on Empirical Methods in Natural Language Processing2025-07-11alphaXiv
arXiv
S2
-
NeoBabel: A Multilingual Open Tower for Visual Generation
-
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
Annual Meeting of the Association for Computational Linguistics2025-06-12alphaXiv
arXiv
S2
-
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
Conference on Empirical Methods in Natural Language Processing2025-05-30alphaXiv
arXiv
S2
-
The Multilingual Divide and Its Impact on Global AI Safety
-
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
-
Aya Vision: Advancing the Frontier of Multilingual Multimodality
-
The Leaderboard Illusion
-
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
-
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
-
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
-
Command A: An Enterprise-Ready Large Language Model
-
From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
Annual Meeting of the Association for Computational Linguistics2025-02-19alphaXiv
arXiv
S2
-
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
North American Chapter of the Association for Computational Linguistics2025-02-04alphaXiv
arXiv
S2
-
Towards Best Practices for Open Datasets for LLM Training
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Annual Meeting of the Association for Computational Linguistics2025S2
-
Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Conference on Machine Translation2025S2
-
RLHF Algorithms Ranked: An Extensive Evaluation Across Diverse Tasks, Rewards, and Hyperparameters
Conference on Empirical Methods in Natural Language Processing2025S2
-
To Code or Not To Code? Exploring Impact of Code in Pre-training
International Conference on Learning Representations2025S2
-
Command-A-Translate: Raising the Bar of Machine Translation with Difficulty Filtering
Conference on Machine Translation2025S2
-
Generating Complex Question Decompositions in the Face of Distribution Shifts
North American Chapter of the Association for Computational Linguistics2025S2
-
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
-
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
International Conference on Learning Representations2024-11-29alphaXiv
arXiv
S2
-
M-RewardBench: Evaluating Reward Models in Multilingual Settings
Annual Meeting of the Association for Computational Linguistics2024-10-20alphaXiv
arXiv
S2
-
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
-
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
-
Mathematical modeling of free vibration of star-shaped auxetic rectangular plate
Archive of applied mechanics (1991)2024-08-21S2
-
To Code, or Not To Code? Exploring Impact of Code in Pre-training
-
LLM See, LLM Do: Leveraging Active Inheritance to Target Non-Differentiable Objectives
Conference on Empirical Methods in Natural Language Processing2024S2
-
Cross-lingual Transfer Dynamics in BLOOMZ: Insights into Multilingual Generalization
-
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification