← People
Nathan Godey

Nathan Godey interacted with

Cornell University
@nthngdyLinkedInGoogle ScholarWebsitePapers in the feed →

Papers · 15
  1. Co-LMLM: Continuous-Query Limited Memory Language Models
    2026-07-08alphaXiv arXiv S2
  2. The State-Prediction Separation Hypothesis
    2026-07-01alphaXiv arXiv S2
  3. Lost in Backpropagation: The LM Head is a Gradient Bottleneck
    arXiv.org2026-03-10alphaXiv arXiv S2
  4. Gaperon: A Peppered English-French Generative Language Model Suite
    Annual Meeting of the Association for Computational Linguistics2026S2
  5. Biomed-Enriched: Data-Efficient Biomedical Pretraining via Paragraph-Level Annotation
    Annual Meeting of the Association for Computational Linguistics2026S2
  6. Gaperon: A Peppered English-French Generative Language Model Suite
    arXiv.org2025-10-29alphaXiv arXiv S2
  7. Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
    arXiv.org2025-06-25alphaXiv arXiv S2
  8. Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
    2025-03-04alphaXiv arXiv S2
  9. Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
    arXiv.org2025S2
  10. Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
    arXiv.org2024-04-11alphaXiv arXiv S2
  11. On the Scaling Laws of Geographical Representation in Language Models
    International Conference on Language Resources and Evaluation2024-02-29alphaXiv arXiv S2
  12. Anisotropy Is Inherent to Self-Attention in Transformers
    Conference of the European Chapter of the Association for Computational Linguistics2024-01-22alphaXiv arXiv S2
  13. Headless Language Models: Learning without Predicting with Contrastive Weight Tying
    International Conference on Learning Representations2023-09-15alphaXiv arXiv S2
  14. Is Anisotropy Inherent to Transformers?
    arXiv.org2023-06-13alphaXiv arXiv S2
  15. MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling
    arXiv.org2022-12-14alphaXiv arXiv S2