← People
Antonio Orvieto
following
MPI / ELLIS Tübingen
@orvieto_antonio
Papers in the feed →
Papers · 46
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
2026-08-06
alphaXiv
arXiv
S2
How the Hessian-Spectrum of Neural Networks Depends on Data
2026-07-15
alphaXiv
arXiv
S2
Muown Implicitly Performs Angular Step-size Decay
arXiv.org
2026-06-22
alphaXiv
arXiv
S2
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
arXiv.org
2026-06-16
alphaXiv
arXiv
S2
Beyond a Single Explanation of the Adam-SGD Gap
arXiv.org
2026-06-12
alphaXiv
arXiv
S2
Towards Understanding Self-Pretraining for Sequence Classification
arXiv.org
2026-05-20
alphaXiv
arXiv
S2
GRASP: Deterministic argument ranking in interaction graphs
arXiv.org
2026-05-18
alphaXiv
arXiv
S2
Muown: Row-Norm Control for Muon Optimization
arXiv.org
2026-05-11
alphaXiv
arXiv
S2
Sequence Modeling Architectures: Foundations [Special Issue on the Mathematics of Deep Learning]
IEEE Signal Processing Magazine
2026-05-01
S2
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
arXiv.org
2026-03-16
alphaXiv
arXiv
S2
GASP: Guided Asymmetric Self-Play For Coding LLMs
arXiv.org
2026-03-16
alphaXiv
arXiv
S2
Improved state mixing in higher-order and block diagonal linear recurrent networks
arXiv.org
2026-02-12
alphaXiv
arXiv
S2
Explaining Grokking in Transformers through the Lens of Inductive Bias
arXiv.org
2026-02-06
alphaXiv
arXiv
S2
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
arXiv.org
2026-01-13
alphaXiv
arXiv
S2
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
Annual Meeting of the Association for Computational Linguistics
2026
S2
Are Current AI Systems Unlocking Knowledge Discovery in Genomics?
Daedalus
2026
S2
Scaling Behavior of Discrete Diffusion Language Models
arXiv.org
2025-12-11
alphaXiv
arXiv
S2
Adam Simplified: Bias Correction Debunked
arXiv.org
2025-11-25
alphaXiv
arXiv
S2
Selective Rotary Position Embedding
arXiv.org
2025-11-21
alphaXiv
arXiv
S2
Design Principles for Sequence Models via Coefficient Dynamics
arXiv.org
2025-10-10
alphaXiv
arXiv
S2
How does the optimizer implicitly bias the model merging loss landscape?
arXiv.org
2025-10-06
alphaXiv
arXiv
S2
Revisiting associative recall in modern recurrent models
2025-08-26
alphaXiv
arXiv
S2
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
Advances in Neural Information Processing Systems 38
2025-08-20
alphaXiv
arXiv
S2
GitChameleon: Evaluating AI Code Generation Against Python Library Version Incompatibilities
2025-07-16
alphaXiv
arXiv
S2
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
arXiv.org
2025-07-15
alphaXiv
arXiv
S2
(Almost) Free Modality Stitching of Foundation Models
Conference on Empirical Methods in Natural Language Processing
2025-07-14
alphaXiv
arXiv
S2
Generalized Linear Mode Connectivity for Transformers
Neural Information Processing Systems
2025-06-28
alphaXiv
arXiv
S2
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
arXiv.org
2025-06-14
alphaXiv
arXiv
S2
On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
2025-05-30
alphaXiv
arXiv
S2
In Search of Adam's Secret Sauce
Neural Information Processing Systems
2025-05-27
alphaXiv
arXiv
S2
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
arXiv.org
2025-04-08
alphaXiv
arXiv
S2
Fixed-Point RNNs: Interpolating from Diagonal to Dense
Neural Information Processing Systems
2025-03-13
alphaXiv
arXiv
S2
Generalized Interpolating Discrete Diffusion
International Conference on Machine Learning
2025-03-06
alphaXiv
arXiv
S2
An Uncertainty Principle for Linear Recurrent Neural Networks
Annual Conference Computational Learning Theory
2025-02-13
alphaXiv
arXiv
S2
When, Where and Why to Average Weights?
International Conference on Machine Learning
2025-02-10
alphaXiv
arXiv
S2
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
International Conference on Learning Representations
2025
S2
Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations
arXiv.org
2025
S2
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
arXiv.org
2024-11-24
alphaXiv
arXiv
S2
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
arXiv.org
2024-10-31
alphaXiv
arXiv
S2
Loss Landscape Characterization of Neural Networks without Over-Parametrization
Neural Information Processing Systems
2024-10-16
alphaXiv
arXiv
S2
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
International Conference on Learning Representations
2024-10-15
alphaXiv
arXiv
S2
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
International Conference on Machine Learning
2024
S2
Why do Learning Rates Transfer? Reconciling Optimization and Scaling Limits for Deep Learning
arXiv.org
2024
S2
Recurrent Distance-Encoding Neural Networks for Graph Representation Learning
arXiv.org
2023
S2
Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning (Appendix)
2023
S2
Escaping Random Teacher Initialization Enhances Signal Propagation and Representation
S2