← People
Catherine Arnett

Catherine Arnett interacted with

EleutherAI
met · ACL 2026 · 2026-07-02
@@linguist_catLinkedInGoogle ScholarWebsitePapers in the feed →

Papers · 20
  1. Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
    arXiv.org2026-06-03alphaXiv arXiv S2
  2. Weight Tying Biases Token Embeddings Towards the Output Space
    Annual Meeting of the Association for Computational Linguistics2026-03-27alphaXiv arXiv S2
  3. How Open Must Language Models be to Enable Reliable Scientific Inference?
    arXiv.org2026-03-27alphaXiv arXiv S2
  4. CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
    Annual Meeting of the Association for Computational Linguistics2026-01-25alphaXiv arXiv S2
  5. Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
    arXiv.org2025-10-28alphaXiv arXiv S2
  6. Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
    2025-10-28alphaXiv arXiv S2
  7. Explaining and Mitigating Crosslingual Tokenizer Inequities
    Neural Information Processing Systems2025-10-24alphaXiv arXiv S2
  8. Evaluating Morphological Alignment of Tokenizers in 70 Languages
    arXiv.org2025-07-08alphaXiv arXiv S2
  9. On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
    Annual Meeting of the Association for Computational Linguistics2025-03-05alphaXiv arXiv S2
  10. Why do language models perform worse for morphologically complex languages?
    International Conference on Computational Linguistics2025S2
  11. Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
    arXiv.org2025S2
  12. Why do language models perform worse for morphologically complex languages?
    arXiv.org2024-11-21alphaXiv arXiv S2
  13. Syntax drives default language selection in bilingual connected speech production
    Journal of Experimental Psychology. Learning, Memory and Cognition2024-10-17S2
  14. Goldfish: Monolingual Language Models for 350 Languages
    arXiv.org2024-08-19alphaXiv arXiv S2
  15. Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
    arXiv.org2024-04-30alphaXiv arXiv S2
  16. Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
    Special Interest Group on Computational Morphology and Phonology Workshop2024-03-20alphaXiv arXiv S2
  17. A Bit of a Problem: Measurement Disparities in Dataset Sizes across Languages
    SIGUL2024-03-01alphaXiv arXiv S2
  18. When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages
    Conference on Empirical Methods in Natural Language Processing2023-11-15alphaXiv arXiv S2
  19. Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models
    Conference on Empirical Methods in Natural Language Processing2023-11-15alphaXiv arXiv S2
  20. Crosslingual Structural Priming and the Pre-Training Dynamics of Bilingual Language Models
    arXiv.org2023-10-11alphaXiv arXiv S2