← People
Leshem Choshen
interacted with
MIT-IBM → new PI at Weizmann
met · ACL 2026 · 2026-07-02
@LChoshen
LinkedIn
Google Scholar
Papers in the feed →
Papers · 121
Automated Discovery Has No Universally Superior Harness
2026-07-20
alphaXiv
arXiv
S2
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
Annual Meeting of the Association for Computational Linguistics
2026-07-09
alphaXiv
arXiv
S2
Cross-Lingual Exploration for Parametric Knowledge
arXiv.org
2026-06-23
alphaXiv
arXiv
S2
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
arXiv.org
2026-06-12
alphaXiv
arXiv
S2
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
arXiv.org
2026-06-08
alphaXiv
arXiv
S2
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
arXiv.org
2026-05-28
alphaXiv
arXiv
S2
Instructions Shape Production of Language, not Processing
arXiv.org
2026-05-11
alphaXiv
arXiv
S2
Mediocrity is the key for LLM as a Judge Anchor Selection
Annual Meeting of the Association for Computational Linguistics
2026-03-17
alphaXiv
arXiv
S2
CUBE: A Standard for Unifying Agent Benchmarks
arXiv.org
2026-03-16
alphaXiv
arXiv
S2
Resolving Interference (RI): Disentangling Models for Improved Model Merging
arXiv.org
2026-03-13
alphaXiv
arXiv
S2
Do LLMs Benefit From Their Own Words?
arXiv.org
2026-02-27
alphaXiv
arXiv
S2
General Agent Evaluation
arXiv.org
2026-02-26
alphaXiv
arXiv
S2
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
arXiv.org
2026-02-23
alphaXiv
arXiv
S2
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
arXiv.org
2026-02-18
alphaXiv
arXiv
S2
Robustness as an Emergent Property of Task Performance
arXiv.org
2026-02-03
alphaXiv
arXiv
S2
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Annual Meeting of the Association for Computational Linguistics
2026-01-25
alphaXiv
arXiv
S2
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
arXiv.org
2026-01-22
alphaXiv
arXiv
S2
Will it Merge? On The Causes of Model Mergeability
Annual Meeting of the Association for Computational Linguistics
2026-01-10
alphaXiv
arXiv
S2
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness in LLMs
Proceedings of the First Workshop on Structured Understanding, Retrieval, and Generation in the LLM Era (SURGeLLM 2026)
2026
S2
A Latent Variable Framework for Scaling Laws in Large Language Models
arXiv.org
2025-12-06
alphaXiv
arXiv
S2
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
2025-10-28
alphaXiv
arXiv
S2
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
Conference of the European Chapter of the Association for Computational Linguistics
2025-10-11
alphaXiv
arXiv
S2
Bigger is not always better: The importance of human-scale language modeling for psycholinguistics
Journal of Memory and Language
2025-10-01
S2
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
arXiv.org
2025-07-22
alphaXiv
arXiv
S2
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Conference on Empirical Methods in Natural Language Processing
2025-07-11
alphaXiv
arXiv
S2
CRISP: Complex Reasoning with Interpretable Step-based Plans
arXiv.org
2025-07-09
alphaXiv
arXiv
S2
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
2025-07-03
alphaXiv
arXiv
S2
Can Gradient Descent Simulate Prompting?
arXiv.org
2025-06-25
alphaXiv
arXiv
S2
TextArena
arXiv.org
2025-04-15
alphaXiv
arXiv
S2
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning
2025-04-10
alphaXiv
arXiv
S2
Pretraining Language Models for Diachronic Linguistic Change Discovery
Conference of the European Chapter of the Association for Computational Linguistics
2025-04-07
alphaXiv
arXiv
S2
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
arXiv.org
2025-03-13
alphaXiv
arXiv
S2
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
Annual Meeting of the Association for Computational Linguistics
2025-03-03
alphaXiv
arXiv
S2
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
arXiv.org
2025-02-26
alphaXiv
arXiv
S2
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
arXiv.org
2025-02-15
alphaXiv
arXiv
S2
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Annual Meeting of the Association for Computational Linguistics
2025
S2
Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
arXiv.org
2025
S2
Findings of the Third BabyLM Challenge: Accelerating Language Modeling Research with Cognitively Plausible Data
Proceedings of the First BabyLM Workshop
2025
S2
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
Trans. Mach. Learn. Res.
2025
S2
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
arXiv.org
2025
S2
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
Neural Information Processing Systems
2024-12-09
alphaXiv
arXiv
S2
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
arXiv.org
2024-12-06
alphaXiv
arXiv
S2
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
arXiv.org
2024-12-04
alphaXiv
arXiv
S2
Holmes ⌕ A Benchmark to Assess the Linguistic Competence of Language Models
Transactions of the Association for Computational Linguistics
2024-12-01
S2
ZipNN: Lossless Compression for AI Models
IEEE International Conference on Cloud Computing
2024-11-07
alphaXiv
arXiv
S2
Model merging with SVD to tie the Knots
International Conference on Learning Representations
2024-10-25
alphaXiv
arXiv
S2
A Hitchhiker's Guide to Scaling Law Estimation
International Conference on Machine Learning
2024-10-15
alphaXiv
arXiv
S2
LiveXiv - A Multi-Modal Live Benchmark Based on Arxiv Papers Content
International Conference on Learning Representations
2024-10-14
alphaXiv
arXiv
S2
Unforgettable Generalization in Language Models
arXiv.org
2024-09-03
alphaXiv
arXiv
S2
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
2024-08-22
alphaXiv
arXiv
S2
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
arXiv.org
2024-08-20
alphaXiv
arXiv
S2
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
Annual Meeting of the Association for Computational Linguistics
2024-08-15
alphaXiv
arXiv
S2
The future of open human feedback
Nature Machine Intelligence
2024-08-15
alphaXiv
arXiv
S2
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
arXiv.org
2024-08-13
alphaXiv
arXiv
S2
Data Contamination Report from the 2024 CONDA Shared Task
CONDA
2024-07-31
alphaXiv
arXiv
S2
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
2024-07-18
alphaXiv
arXiv
S2
Naturally Occurring Feedback is Common, Extractable and Useful
2024-07-15
alphaXiv
arXiv
S2
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
International Conference on Machine Learning
2024-06-17
alphaXiv
arXiv
S2
Efficient multi-prompt evaluation of LLMs
Neural Information Processing Systems
2024-05-27
alphaXiv
arXiv
S2
Elements of World Knowledge (EWOK): A cognition-inspired framework for evaluating basic world knowledge in language models
Transactions of the Association for Computational Linguistics
2024-05-15
alphaXiv
arXiv
S2
Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs)
International Conference on Language Resources and Evaluation
2024-05-01
S2
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
2024-04-29
alphaXiv
arXiv
S2
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
arXiv.org
2024-04-09
alphaXiv
arXiv
S2
Lossless and Near-Lossless Compression for Foundation Models
arXiv.org
2024-04-05
alphaXiv
arXiv
S2
NumeroLogic: Number Encoding for Enhanced LLMs’ Numerical Reasoning
Conference on Empirical Methods in Natural Language Processing
2024-03-30
alphaXiv
arXiv
S2
Asymmetry in Low-Rank Adapters of Foundation Models
International Conference on Machine Learning
2024-02-26
alphaXiv
arXiv
S2
tinyBenchmarks: evaluating LLMs with fewer examples
International Conference on Machine Learning
2024-02-22
alphaXiv
arXiv
S2
Label-Efficient Model Selection for Text Generation
Annual Meeting of the Association for Computational Linguistics
2024-02-12
alphaXiv
arXiv
S2
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
North American Chapter of the Association for Computational Linguistics
2024-01-25
alphaXiv
arXiv
S2
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
arXiv.org
2024-01-25
alphaXiv
arXiv
S2
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
Annual Meeting of the Association for Computational Linguistics
2024-01-16
alphaXiv
arXiv
S2
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
Trans. Mach. Learn. Res.
2023-11-22
alphaXiv
arXiv
S2
Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney
Conference on Empirical Methods in Natural Language Processing
2023-11-20
alphaXiv
arXiv
S2
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
Conference on Empirical Methods in Natural Language Processing
2023-11-13
alphaXiv
arXiv
S2
Efficient Benchmarking (of Language Models)
North American Chapter of the Association for Computational Linguistics
2023-08-22
alphaXiv
arXiv
S2
TIES-Merging: Resolving Interference When Merging Models
Neural Information Processing Systems
2023-06-02
alphaXiv
arXiv
S2
MuLER: Detailed and Scalable Reference-based Evaluation
Conference on Computational Natural Language Learning
2023-05-24
alphaXiv
arXiv
S2
Jump to Conclusions: Short-Cutting Transformers with Linear Transformations
International Conference on Language Resources and Evaluation
2023-03-16
alphaXiv
arXiv
S2
Knowledge is a Region in Weight Space for Fine-tuned Language Models
Conference on Empirical Methods in Natural Language Processing
2023-02-09
alphaXiv
arXiv
S2
Call for Papers - The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
arXiv.org
2023-01-27
alphaXiv
arXiv
S2
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning
Annual Meeting of the Association for Computational Linguistics
2022-12-02
alphaXiv
arXiv
S2
DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
Annual Meeting of the Association for Computational Linguistics
2022-11-10
alphaXiv
arXiv
S2
Where to start? Analyzing the potential value of intermediate models
Conference on Empirical Methods in Natural Language Processing
2022-10-31
alphaXiv
arXiv
S2
Reinforcement Learning with Large Action Spaces for Neural Machine Translation
International Conference on Computational Linguistics
2022-10-06
alphaXiv
arXiv
S2
Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours
Conference on Empirical Methods in Natural Language Processing
2022-08-02
alphaXiv
arXiv
S2
PreQuEL: Quality Estimation of Machine Translation Outputs in Advance
Conference on Empirical Methods in Natural Language Processing
2022-05-18
alphaXiv
arXiv
S2
Some Grammatical Errors are Frequent, Others are Important
arXiv.org
2022-05-11
alphaXiv
arXiv
S2
Fusing finetuned models for better pretraining
arXiv.org
2022-04-06
alphaXiv
arXiv
S2
Cluster & Tune: Boost Cold Start Performance in Text Classification
Annual Meeting of the Association for Computational Linguistics
2022-03-20
alphaXiv
arXiv
S2
Semantics-aware Attention Improves Neural Machine Translation
STARSEM
2021-10-13
alphaXiv
arXiv
S2
On Neurons Invariant to Sentence Structural Changes in Neural Machine Translation
Conference on Computational Natural Language Learning
2021-10-06
alphaXiv
arXiv
S2
The Grammar-Learning Trajectories of Neural Language Models
Annual Meeting of the Association for Computational Linguistics
2021-09-13
alphaXiv
arXiv
S2
ComSum: Commit Messages Summarization and Meaning Preservation
arXiv.org
2021-08-23
alphaXiv
arXiv
S2
Part of Speech and Universal Dependency effects on English Arabic Machine Translation
arXiv.org
2021-06-01
alphaXiv
arXiv
S2
Q^{2}: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
Conference on Empirical Methods in Natural Language Processing
2021-04-16
alphaXiv
arXiv
S2
Mediators in Determining what Processing BERT Performs First
North American Chapter of the Association for Computational Linguistics
2021-04-13
alphaXiv
arXiv
S2
GrASP: A Library for Extracting and Exploring Human-Interpretable Textual Patterns
International Conference on Language Resources and Evaluation
2021-04-08
alphaXiv
arXiv
S2
SERRANT: a syntactic classifier for English Grammatical Error Types
arXiv.org
2021-04-06
alphaXiv
arXiv
S2
An autonomous debating system
Nature
2021-03-01
S2
Transition based Graph Decoder for Neural Machine Translation
arXiv.org
2021-01-29
S2
Enhancing the Transformer Decoder with Transition-based Syntax
Conference on Computational Natural Language Learning
2021-01-29
alphaXiv
arXiv
S2
Active Learning for BERT: An Empirical Study
Conference on Empirical Methods in Natural Language Processing
2020-11-01
S2
Classifying Syntactic Errors in Learner Language
Conference on Computational Natural Language Learning
2020-10-21
alphaXiv
arXiv
S2
Unsupervised Expressive Rules Provide Explainability and Assist Human Experts Grasping New Domains
Findings
2020-10-19
alphaXiv
arXiv
S2
Corpus Wide Argument Mining - a Working Solution
AAAI Conference on Artificial Intelligence
2019-11-25
alphaXiv
arXiv
S2
All Neural Networks are Created Equal
arXiv.org
2019-09-25
S2
Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation
Conference on Computational Natural Language Learning
2019-09-15
alphaXiv
arXiv
S2
On the Weaknesses of Reinforcement Learning for Neural Machine Translation
International Conference on Learning Representations
2019-07-03
alphaXiv
arXiv
S2
Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network
Annual Meeting of the Association for Computational Linguistics
2019-07-01
alphaXiv
arXiv
S2
Learning to combine Grammatical Error Corrections
BEA@ACL
2019-06-10
alphaXiv
arXiv
S2
Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets
International Conference on Machine Learning
2019-05-26
alphaXiv
arXiv
S2
The Language of Legal and Illegal Activity on the Darknet
Annual Meeting of the Association for Computational Linguistics
2019-05-14
alphaXiv
arXiv
S2
SemEval-2019 Task 1: Cross-lingual Semantic Parsing with UCCA
International Workshop on Semantic Evaluation
2019-03-06
alphaXiv
arXiv
S2
Will it Blend? Blending Weak and Strong Labeled Data in a Neural Network for Argumentation Mining
Annual Meeting of the Association for Computational Linguistics
2018-07-01
S2
SemEval 2019 Shared Task: Cross-lingual Semantic Parsing with UCCA - Call for Participation
arXiv.org
2018-05-31
alphaXiv
arXiv
S2
Inherent Biases in Reference-based Evaluation for Grammatical Error Correction
Annual Meeting of the Association for Computational Linguistics
2018-04-30
alphaXiv
arXiv
S2
Automatic Metric Validation for Grammatical Error Correction
Annual Meeting of the Association for Computational Linguistics
2018-04-30
alphaXiv
arXiv
S2
Reference-less Measure of Faithfulness for Grammatical Error Correction
North American Chapter of the Association for Computational Linguistics
2018-04-11
alphaXiv
arXiv
S2
DORA The Explorer: Directed Outreaching Reinforcement Action-Selection
International Conference on Learning Representations
2018-02-15
alphaXiv
arXiv
S2
Journal of Memory and Language
S2
P OSITION : A GENTIC S YSTEMS S HOULD BE G ENERAL
S2