← People
Antonio Orvieto

Antonio Orvieto following

MPI / ELLIS Tübingen
@orvieto_antonioPapers in the feed →

Papers · 46
  1. Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
    2026-08-06alphaXiv arXiv S2
  2. How the Hessian-Spectrum of Neural Networks Depends on Data
    2026-07-15alphaXiv arXiv S2
  3. Muown Implicitly Performs Angular Step-size Decay
    arXiv.org2026-06-22alphaXiv arXiv S2
  4. Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
    arXiv.org2026-06-16alphaXiv arXiv S2
  5. Beyond a Single Explanation of the Adam-SGD Gap
    arXiv.org2026-06-12alphaXiv arXiv S2
  6. Towards Understanding Self-Pretraining for Sequence Classification
    arXiv.org2026-05-20alphaXiv arXiv S2
  7. GRASP: Deterministic argument ranking in interaction graphs
    arXiv.org2026-05-18alphaXiv arXiv S2
  8. Muown: Row-Norm Control for Muon Optimization
    arXiv.org2026-05-11alphaXiv arXiv S2
  9. Sequence Modeling Architectures: Foundations [Special Issue on the Mathematics of Deep Learning]
    IEEE Signal Processing Magazine2026-05-01S2
  10. Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
    arXiv.org2026-03-16alphaXiv arXiv S2
  11. GASP: Guided Asymmetric Self-Play For Coding LLMs
    arXiv.org2026-03-16alphaXiv arXiv S2
  12. Improved state mixing in higher-order and block diagonal linear recurrent networks
    arXiv.org2026-02-12alphaXiv arXiv S2
  13. Explaining Grokking in Transformers through the Lens of Inductive Bias
    arXiv.org2026-02-06alphaXiv arXiv S2
  14. Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
    arXiv.org2026-01-13alphaXiv arXiv S2
  15. GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
    Annual Meeting of the Association for Computational Linguistics2026S2
  16. Are Current AI Systems Unlocking Knowledge Discovery in Genomics?
    Daedalus2026S2
  17. Scaling Behavior of Discrete Diffusion Language Models
    arXiv.org2025-12-11alphaXiv arXiv S2
  18. Adam Simplified: Bias Correction Debunked
    arXiv.org2025-11-25alphaXiv arXiv S2
  19. Selective Rotary Position Embedding
    arXiv.org2025-11-21alphaXiv arXiv S2
  20. Design Principles for Sequence Models via Coefficient Dynamics
    arXiv.org2025-10-10alphaXiv arXiv S2
  21. How does the optimizer implicitly bias the model merging loss landscape?
    arXiv.org2025-10-06alphaXiv arXiv S2
  22. Revisiting associative recall in modern recurrent models
    2025-08-26alphaXiv arXiv S2
  23. Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
    Advances in Neural Information Processing Systems 382025-08-20alphaXiv arXiv S2
  24. GitChameleon: Evaluating AI Code Generation Against Python Library Version Incompatibilities
    2025-07-16alphaXiv arXiv S2
  25. Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
    arXiv.org2025-07-15alphaXiv arXiv S2
  26. (Almost) Free Modality Stitching of Foundation Models
    Conference on Empirical Methods in Natural Language Processing2025-07-14alphaXiv arXiv S2
  27. Generalized Linear Mode Connectivity for Transformers
    Neural Information Processing Systems2025-06-28alphaXiv arXiv S2
  28. Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
    arXiv.org2025-06-14alphaXiv arXiv S2
  29. On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
    2025-05-30alphaXiv arXiv S2
  30. In Search of Adam's Secret Sauce
    Neural Information Processing Systems2025-05-27alphaXiv arXiv S2
  31. Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
    arXiv.org2025-04-08alphaXiv arXiv S2
  32. Fixed-Point RNNs: Interpolating from Diagonal to Dense
    Neural Information Processing Systems2025-03-13alphaXiv arXiv S2
  33. Generalized Interpolating Discrete Diffusion
    International Conference on Machine Learning2025-03-06alphaXiv arXiv S2
  34. An Uncertainty Principle for Linear Recurrent Neural Networks
    Annual Conference Computational Learning Theory2025-02-13alphaXiv arXiv S2
  35. When, Where and Why to Average Weights?
    International Conference on Machine Learning2025-02-10alphaXiv arXiv S2
  36. Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
    International Conference on Learning Representations2025S2
  37. Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations
    arXiv.org2025S2
  38. Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
    arXiv.org2024-11-24alphaXiv arXiv S2
  39. NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
    arXiv.org2024-10-31alphaXiv arXiv S2
  40. Loss Landscape Characterization of Neural Networks without Over-Parametrization
    Neural Information Processing Systems2024-10-16alphaXiv arXiv S2
  41. Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
    International Conference on Learning Representations2024-10-15alphaXiv arXiv S2
  42. Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
    International Conference on Machine Learning2024S2
  43. Why do Learning Rates Transfer? Reconciling Optimization and Scaling Limits for Deep Learning
    arXiv.org2024S2
  44. Recurrent Distance-Encoding Neural Networks for Graph Representation Learning
    arXiv.org2023S2
  45. Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning (Appendix)
    2023S2
  46. Escaping Random Teacher Initialization Enhances Signal Propagation and Representation
    S2