← People
Leshem Choshen

Leshem Choshen interacted with

MIT-IBM → new PI at Weizmann
met · ACL 2026 · 2026-07-02
@LChoshenLinkedInGoogle ScholarPapers in the feed →

Papers · 121
  1. Automated Discovery Has No Universally Superior Harness
    2026-07-20alphaXiv arXiv S2
  2. Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
    Annual Meeting of the Association for Computational Linguistics2026-07-09alphaXiv arXiv S2
  3. Cross-Lingual Exploration for Parametric Knowledge
    arXiv.org2026-06-23alphaXiv arXiv S2
  4. Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
    arXiv.org2026-06-12alphaXiv arXiv S2
  5. Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
    arXiv.org2026-06-08alphaXiv arXiv S2
  6. MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
    arXiv.org2026-05-28alphaXiv arXiv S2
  7. Instructions Shape Production of Language, not Processing
    arXiv.org2026-05-11alphaXiv arXiv S2
  8. Mediocrity is the key for LLM as a Judge Anchor Selection
    Annual Meeting of the Association for Computational Linguistics2026-03-17alphaXiv arXiv S2
  9. CUBE: A Standard for Unifying Agent Benchmarks
    arXiv.org2026-03-16alphaXiv arXiv S2
  10. Resolving Interference (RI): Disentangling Models for Improved Model Merging
    arXiv.org2026-03-13alphaXiv arXiv S2
  11. Do LLMs Benefit From Their Own Words?
    arXiv.org2026-02-27alphaXiv arXiv S2
  12. General Agent Evaluation
    arXiv.org2026-02-26alphaXiv arXiv S2
  13. BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
    arXiv.org2026-02-23alphaXiv arXiv S2
  14. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
    arXiv.org2026-02-18alphaXiv arXiv S2
  15. Robustness as an Emergent Property of Task Performance
    arXiv.org2026-02-03alphaXiv arXiv S2
  16. CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
    Annual Meeting of the Association for Computational Linguistics2026-01-25alphaXiv arXiv S2
  17. ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
    arXiv.org2026-01-22alphaXiv arXiv S2
  18. Will it Merge? On The Causes of Model Mergeability
    Annual Meeting of the Association for Computational Linguistics2026-01-10alphaXiv arXiv S2
  19. The Mighty ToRR: A Benchmark for Table Reasoning and Robustness in LLMs
    Proceedings of the First Workshop on Structured Understanding, Retrieval, and Generation in the LLM Era (SURGeLLM 2026)2026S2
  20. A Latent Variable Framework for Scaling Laws in Large Language Models
    arXiv.org2025-12-06alphaXiv arXiv S2
  21. Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
    2025-10-28alphaXiv arXiv S2
  22. BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
    Conference of the European Chapter of the Association for Computational Linguistics2025-10-11alphaXiv arXiv S2
  23. Bigger is not always better: The importance of human-scale language modeling for psycholinguistics
    Journal of Memory and Language2025-10-01S2
  24. Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
    arXiv.org2025-07-22alphaXiv arXiv S2
  25. From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation
    Conference on Empirical Methods in Natural Language Processing2025-07-11alphaXiv arXiv S2
  26. CRISP: Complex Reasoning with Interpretable Step-based Plans
    arXiv.org2025-07-09alphaXiv arXiv S2
  27. LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
    2025-07-03alphaXiv arXiv S2
  28. Can Gradient Descent Simulate Prompting?
    arXiv.org2025-06-25alphaXiv arXiv S2
  29. TextArena
    arXiv.org2025-04-15alphaXiv arXiv S2
  30. Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
    Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning2025-04-10alphaXiv arXiv S2
  31. Pretraining Language Models for Diachronic Linguistic Change Discovery
    Conference of the European Chapter of the Association for Computational Linguistics2025-04-07alphaXiv arXiv S2
  32. NeurIPS 2023 LLM Efficiency Fine-tuning Competition
    arXiv.org2025-03-13alphaXiv arXiv S2
  33. DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
    Annual Meeting of the Association for Computational Linguistics2025-03-03alphaXiv arXiv S2
  34. The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
    arXiv.org2025-02-26alphaXiv arXiv S2
  35. BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
    arXiv.org2025-02-15alphaXiv arXiv S2
  36. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    Annual Meeting of the Association for Computational Linguistics2025S2
  37. Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
    arXiv.org2025S2
  38. Findings of the Third BabyLM Challenge: Accelerating Language Modeling Research with Cognitively Plausible Data
    Proceedings of the First BabyLM Workshop2025S2
  39. A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
    Trans. Mach. Learn. Res.2025S2
  40. LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
    arXiv.org2025S2
  41. Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
    Neural Information Processing Systems2024-12-09alphaXiv arXiv S2
  42. Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
    arXiv.org2024-12-06alphaXiv arXiv S2
  43. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
    arXiv.org2024-12-04alphaXiv arXiv S2
  44. Holmes ⌕ A Benchmark to Assess the Linguistic Competence of Language Models
    Transactions of the Association for Computational Linguistics2024-12-01S2
  45. ZipNN: Lossless Compression for AI Models
    IEEE International Conference on Cloud Computing2024-11-07alphaXiv arXiv S2
  46. Model merging with SVD to tie the Knots
    International Conference on Learning Representations2024-10-25alphaXiv arXiv S2
  47. A Hitchhiker's Guide to Scaling Law Estimation
    International Conference on Machine Learning2024-10-15alphaXiv arXiv S2
  48. LiveXiv - A Multi-Modal Live Benchmark Based on Arxiv Papers Content
    International Conference on Learning Representations2024-10-14alphaXiv arXiv S2
  49. Unforgettable Generalization in Language Models
    arXiv.org2024-09-03alphaXiv arXiv S2
  50. How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
    2024-08-22alphaXiv arXiv S2
  51. Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
    arXiv.org2024-08-20alphaXiv arXiv S2
  52. The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
    Annual Meeting of the Association for Computational Linguistics2024-08-15alphaXiv arXiv S2
  53. The future of open human feedback
    Nature Machine Intelligence2024-08-15alphaXiv arXiv S2
  54. A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
    arXiv.org2024-08-13alphaXiv arXiv S2
  55. Data Contamination Report from the 2024 CONDA Shared Task
    CONDA2024-07-31alphaXiv arXiv S2
  56. Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
    2024-07-18alphaXiv arXiv S2
  57. Naturally Occurring Feedback is Common, Extractable and Useful
    2024-07-15alphaXiv arXiv S2
  58. Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
    International Conference on Machine Learning2024-06-17alphaXiv arXiv S2
  59. Efficient multi-prompt evaluation of LLMs
    Neural Information Processing Systems2024-05-27alphaXiv arXiv S2
  60. Elements of World Knowledge (EWOK): A cognition-inspired framework for evaluating basic world knowledge in language models
    Transactions of the Association for Computational Linguistics2024-05-15alphaXiv arXiv S2
  61. Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs)
    International Conference on Language Resources and Evaluation2024-05-01S2
  62. Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
    2024-04-29alphaXiv arXiv S2
  63. [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
    arXiv.org2024-04-09alphaXiv arXiv S2
  64. Lossless and Near-Lossless Compression for Foundation Models
    arXiv.org2024-04-05alphaXiv arXiv S2
  65. NumeroLogic: Number Encoding for Enhanced LLMs’ Numerical Reasoning
    Conference on Empirical Methods in Natural Language Processing2024-03-30alphaXiv arXiv S2
  66. Asymmetry in Low-Rank Adapters of Foundation Models
    International Conference on Machine Learning2024-02-26alphaXiv arXiv S2
  67. tinyBenchmarks: evaluating LLMs with fewer examples
    International Conference on Machine Learning2024-02-22alphaXiv arXiv S2
  68. Label-Efficient Model Selection for Text Generation
    Annual Meeting of the Association for Computational Linguistics2024-02-12alphaXiv arXiv S2
  69. Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
    North American Chapter of the Association for Computational Linguistics2024-01-25alphaXiv arXiv S2
  70. Genie: Achieving Human Parity in Content-Grounded Datasets Generation
    arXiv.org2024-01-25alphaXiv arXiv S2
  71. Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
    Annual Meeting of the Association for Computational Linguistics2024-01-16alphaXiv arXiv S2
  72. ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
    Trans. Mach. Learn. Res.2023-11-22alphaXiv arXiv S2
  73. Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney
    Conference on Empirical Methods in Natural Language Processing2023-11-20alphaXiv arXiv S2
  74. Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
    Conference on Empirical Methods in Natural Language Processing2023-11-13alphaXiv arXiv S2
  75. Efficient Benchmarking (of Language Models)
    North American Chapter of the Association for Computational Linguistics2023-08-22alphaXiv arXiv S2
  76. TIES-Merging: Resolving Interference When Merging Models
    Neural Information Processing Systems2023-06-02alphaXiv arXiv S2
  77. MuLER: Detailed and Scalable Reference-based Evaluation
    Conference on Computational Natural Language Learning2023-05-24alphaXiv arXiv S2
  78. Jump to Conclusions: Short-Cutting Transformers with Linear Transformations
    International Conference on Language Resources and Evaluation2023-03-16alphaXiv arXiv S2
  79. Knowledge is a Region in Weight Space for Fine-tuned Language Models
    Conference on Empirical Methods in Natural Language Processing2023-02-09alphaXiv arXiv S2
  80. Call for Papers - The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
    arXiv.org2023-01-27alphaXiv arXiv S2
  81. ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning
    Annual Meeting of the Association for Computational Linguistics2022-12-02alphaXiv arXiv S2
  82. DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
    Annual Meeting of the Association for Computational Linguistics2022-11-10alphaXiv arXiv S2
  83. Where to start? Analyzing the potential value of intermediate models
    Conference on Empirical Methods in Natural Language Processing2022-10-31alphaXiv arXiv S2
  84. Reinforcement Learning with Large Action Spaces for Neural Machine Translation
    International Conference on Computational Linguistics2022-10-06alphaXiv arXiv S2
  85. Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours
    Conference on Empirical Methods in Natural Language Processing2022-08-02alphaXiv arXiv S2
  86. PreQuEL: Quality Estimation of Machine Translation Outputs in Advance
    Conference on Empirical Methods in Natural Language Processing2022-05-18alphaXiv arXiv S2
  87. Some Grammatical Errors are Frequent, Others are Important
    arXiv.org2022-05-11alphaXiv arXiv S2
  88. Fusing finetuned models for better pretraining
    arXiv.org2022-04-06alphaXiv arXiv S2
  89. Cluster & Tune: Boost Cold Start Performance in Text Classification
    Annual Meeting of the Association for Computational Linguistics2022-03-20alphaXiv arXiv S2
  90. Semantics-aware Attention Improves Neural Machine Translation
    STARSEM2021-10-13alphaXiv arXiv S2
  91. On Neurons Invariant to Sentence Structural Changes in Neural Machine Translation
    Conference on Computational Natural Language Learning2021-10-06alphaXiv arXiv S2
  92. The Grammar-Learning Trajectories of Neural Language Models
    Annual Meeting of the Association for Computational Linguistics2021-09-13alphaXiv arXiv S2
  93. ComSum: Commit Messages Summarization and Meaning Preservation
    arXiv.org2021-08-23alphaXiv arXiv S2
  94. Part of Speech and Universal Dependency effects on English Arabic Machine Translation
    arXiv.org2021-06-01alphaXiv arXiv S2
  95. Q^{2}: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
    Conference on Empirical Methods in Natural Language Processing2021-04-16alphaXiv arXiv S2
  96. Mediators in Determining what Processing BERT Performs First
    North American Chapter of the Association for Computational Linguistics2021-04-13alphaXiv arXiv S2
  97. GrASP: A Library for Extracting and Exploring Human-Interpretable Textual Patterns
    International Conference on Language Resources and Evaluation2021-04-08alphaXiv arXiv S2
  98. SERRANT: a syntactic classifier for English Grammatical Error Types
    arXiv.org2021-04-06alphaXiv arXiv S2
  99. An autonomous debating system
    Nature2021-03-01S2
  100. Transition based Graph Decoder for Neural Machine Translation
    arXiv.org2021-01-29S2
  101. Enhancing the Transformer Decoder with Transition-based Syntax
    Conference on Computational Natural Language Learning2021-01-29alphaXiv arXiv S2
  102. Active Learning for BERT: An Empirical Study
    Conference on Empirical Methods in Natural Language Processing2020-11-01S2
  103. Classifying Syntactic Errors in Learner Language
    Conference on Computational Natural Language Learning2020-10-21alphaXiv arXiv S2
  104. Unsupervised Expressive Rules Provide Explainability and Assist Human Experts Grasping New Domains
    Findings2020-10-19alphaXiv arXiv S2
  105. Corpus Wide Argument Mining - a Working Solution
    AAAI Conference on Artificial Intelligence2019-11-25alphaXiv arXiv S2
  106. All Neural Networks are Created Equal
    arXiv.org2019-09-25S2
  107. Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation
    Conference on Computational Natural Language Learning2019-09-15alphaXiv arXiv S2
  108. On the Weaknesses of Reinforcement Learning for Neural Machine Translation
    International Conference on Learning Representations2019-07-03alphaXiv arXiv S2
  109. Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network
    Annual Meeting of the Association for Computational Linguistics2019-07-01alphaXiv arXiv S2
  110. Learning to combine Grammatical Error Corrections
    BEA@ACL2019-06-10alphaXiv arXiv S2
  111. Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets
    International Conference on Machine Learning2019-05-26alphaXiv arXiv S2
  112. The Language of Legal and Illegal Activity on the Darknet
    Annual Meeting of the Association for Computational Linguistics2019-05-14alphaXiv arXiv S2
  113. SemEval-2019 Task 1: Cross-lingual Semantic Parsing with UCCA
    International Workshop on Semantic Evaluation2019-03-06alphaXiv arXiv S2
  114. Will it Blend? Blending Weak and Strong Labeled Data in a Neural Network for Argumentation Mining
    Annual Meeting of the Association for Computational Linguistics2018-07-01S2
  115. SemEval 2019 Shared Task: Cross-lingual Semantic Parsing with UCCA - Call for Participation
    arXiv.org2018-05-31alphaXiv arXiv S2
  116. Inherent Biases in Reference-based Evaluation for Grammatical Error Correction
    Annual Meeting of the Association for Computational Linguistics2018-04-30alphaXiv arXiv S2
  117. Automatic Metric Validation for Grammatical Error Correction
    Annual Meeting of the Association for Computational Linguistics2018-04-30alphaXiv arXiv S2
  118. Reference-less Measure of Faithfulness for Grammatical Error Correction
    North American Chapter of the Association for Computational Linguistics2018-04-11alphaXiv arXiv S2
  119. DORA The Explorer: Directed Outreaching Reinforcement Action-Selection
    International Conference on Learning Representations2018-02-15alphaXiv arXiv S2
  120. Journal of Memory and Language
    S2
  121. P OSITION : A GENTIC S YSTEMS S HOULD BE G ENERAL
    S2