← People
Caglar Gulcehre
following
EPFL
@caglarml
Papers in the feed →
Papers · 35
Context-Aware Toxicity Detection in Game Chat: Domain-Adaptive Pretraining with Match Metadata under Limited Labels
Games
2026-07-17
S2
Diffuse AI Control on Fuzzy Tasks
arXiv.org
2026-06-08
alphaXiv
arXiv
S2
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
arXiv.org
2026-06-01
alphaXiv
arXiv
S2
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
arXiv.org
2026-05-27
alphaXiv
arXiv
S2
The Future of Facts: Tracing the Factual Generation-Verification Gap
arXiv.org
2026-05-26
alphaXiv
arXiv
S2
Language Modeling with Hyperspherical Flows
arXiv.org
2026-05-11
alphaXiv
arXiv
S2
Sequence Modeling Architectures: Foundations [Special Issue on the Mathematics of Deep Learning]
IEEE Signal Processing Magazine
2026-05-01
S2
The Diffusion Duality, Chapter II: Ψ-Samplers and Efficient Curriculum
arXiv.org
2026-02-24
alphaXiv
arXiv
S2
RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference
arXiv.org
2026-02-20
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Annual Meeting of the Association for Computational Linguistics
2026
S2
Learning Vision-Language Alignment in Unified LLMs with 24 Text Tokens per Image
International Workshop on Spoken Language Translation
2026
S2
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
arXiv.org
2025-10-22
alphaXiv
arXiv
S2
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
arXiv.org
2025-10-10
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
2025-09-17
alphaXiv
arXiv
S2
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
Neural Information Processing Systems
2025-07-10
alphaXiv
arXiv
S2
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
Advances in Neural Information Processing Systems 38
2025-07-06
alphaXiv
arXiv
S2
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
arXiv.org
2025-06-11
alphaXiv
arXiv
S2
Control Tax: The Price of Keeping AI in Check
arXiv.org
2025-06-05
alphaXiv
arXiv
S2
Partition Generative Modeling: Masked Modeling Without Masks
arXiv.org
2025-05-24
alphaXiv
arXiv
S2
Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
arXiv.org
2025-04-07
alphaXiv
arXiv
S2
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata
arXiv.org
2025-04-02
alphaXiv
arXiv
S2
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
arXiv.org
2025-02-14
alphaXiv
arXiv
S2
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
International Conference on AI in Finance
2025-02-04
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
arXiv.org
2025
S2
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
Neural Information Processing Systems
2025
S2
RAT: Bridging RNN Efficiency and Attention Accuracy in Language Modeling
arXiv.org
2025
S2
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
2024-10-28
alphaXiv
arXiv
S2
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
International Conference on Learning Representations
2024-10-28
alphaXiv
arXiv
S2
SIKeD: Self-guided Iterative Knowledge Distillation for mathematical reasoning
Annual Meeting of the Association for Computational Linguistics
2024-10-24
alphaXiv
arXiv
S2
The Role of Deep Learning Regularizations on Actors in Offline RL
arXiv.org
2024-09-11
alphaXiv
arXiv
S2
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
International Conference on Machine Learning
2024
S2
The Effect of Scheduling and Preemption on the Efficiency of LLM Inference Serving
arXiv.org
2024
S2
Fleet of Agents: Coordinated Problem Solving with Large Language Models using Genetic Particle Filtering
arXiv.org
2024
S2
Building on Efficient Foundations: Effective Training of LLMs with Structured Feedforward Layers
Advances in Neural Information Processing Systems 37
2024
S2
Building on Efficient Foundations: Effective Training of LLMs with Structured Feedforward Layers
Neural Information Processing Systems
2024
S2