← People
Giorgia Ramponi

Giorgia Ramponi following

UZH.AI Hub, ETH AI Center
WebsitePapers in the feed →

Papers · 20
  1. Multi-agent imitation learning with function approximation: Linear Markov games and beyond
    arXiv.org2026-02-26alphaXiv arXiv S2
  2. Aligning Language Models from User Interactions
    arXiv.org2026-02-18alphaXiv arXiv S2
  3. MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
    arXiv.org2026-02-16alphaXiv arXiv S2
  4. Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
    arXiv.org2026-02-02alphaXiv arXiv S2
  5. From Words To Rewards: Leveraging Natural Language For Reinforcement Learning
    Trans. Mach. Learn. Res.2026S2
  6. Rate optimal learning of equilibria from data
    arXiv.org2025-10-10alphaXiv arXiv S2
  7. Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
    arXiv.org2025-09-30alphaXiv arXiv S2
  8. Learning Acrobatic Flight from Preferences
    2025-08-26alphaXiv arXiv S2
  9. Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
    arXiv.org2025-06-29alphaXiv arXiv S2
  10. Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
    arXiv.org2025-06-08alphaXiv arXiv S2
  11. Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
    Neural Information Processing Systems2025-05-23alphaXiv arXiv S2
  12. Clustered KL-barycenter design for policy evaluation
    arXiv.org2025-03-04alphaXiv arXiv S2
  13. Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
    arXiv.org2025-02-13alphaXiv arXiv S2
  14. Non-rectangular Robust MDPs with Normed Uncertainty Sets
    Neural Information Processing Systems2025S2
  15. On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
    Neural Information Processing Systems2024-11-22alphaXiv arXiv S2
  16. Learning Collusion in Episodic, Inventory-Constrained Markets
    Adaptive Agents and Multi-Agent Systems2024-10-24alphaXiv arXiv S2
  17. On the Convergence of Single-Timescale Actor-Critic
    Advances in Neural Information Processing Systems 382024-10-11alphaXiv arXiv S2
  18. Reinforcement Learning from Human Text Feedback: Learning a Reward Model from Human Text Input
    S2
  19. HIP-RL: Hallucinated Inputs for Preference-based Reinforcement Learning in Continuous Domains
    S2
  20. On Imitation in Mean-field Games
    S2