← People
Nathan Godey
interacted with
Cornell University
@nthngdy
LinkedIn
Google Scholar
Website
Papers in the feed →
Papers · 15
Co-LMLM: Continuous-Query Limited Memory Language Models
2026-07-08
alphaXiv
arXiv
S2
The State-Prediction Separation Hypothesis
2026-07-01
alphaXiv
arXiv
S2
Lost in Backpropagation: The LM Head is a Gradient Bottleneck
arXiv.org
2026-03-10
alphaXiv
arXiv
S2
Gaperon: A Peppered English-French Generative Language Model Suite
Annual Meeting of the Association for Computational Linguistics
2026
S2
Biomed-Enriched: Data-Efficient Biomedical Pretraining via Paragraph-Level Annotation
Annual Meeting of the Association for Computational Linguistics
2026
S2
Gaperon: A Peppered English-French Generative Language Model Suite
arXiv.org
2025-10-29
alphaXiv
arXiv
S2
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
arXiv.org
2025-06-25
alphaXiv
arXiv
S2
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
2025-03-04
alphaXiv
arXiv
S2
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
arXiv.org
2025
S2
Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck
arXiv.org
2024-04-11
alphaXiv
arXiv
S2
On the Scaling Laws of Geographical Representation in Language Models
International Conference on Language Resources and Evaluation
2024-02-29
alphaXiv
arXiv
S2
Anisotropy Is Inherent to Self-Attention in Transformers
Conference of the European Chapter of the Association for Computational Linguistics
2024-01-22
alphaXiv
arXiv
S2
Headless Language Models: Learning without Predicting with Contrastive Weight Tying
International Conference on Learning Representations
2023-09-15
alphaXiv
arXiv
S2
Is Anisotropy Inherent to Transformers?
arXiv.org
2023-06-13
alphaXiv
arXiv
S2
MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling
arXiv.org
2022-12-14
alphaXiv
arXiv
S2