← People
Barna Pásztor
following
ETH Zurich
@pasztorb
Papers in the feed →
Papers · 7
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
arXiv.org
2026-03-10
alphaXiv
arXiv
S2
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
arXiv.org
2026-02-27
alphaXiv
arXiv
S2
Aligning Language Models from User Interactions
arXiv.org
2026-02-18
alphaXiv
arXiv
S2
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Annual Meeting of the Association for Computational Linguistics
2026
S2
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
arXiv.org
2025-12-18
alphaXiv
arXiv
S2
Scalable ride-sourcing vehicle rebalancing with service accessibility guarantee: A constrained mean-field reinforcement learning approach
Transportation Research Part C: Emerging Technologies
2025-03-31
alphaXiv
arXiv
S2
Learning Collusion in Episodic, Inventory-Constrained Markets
Adaptive Agents and Multi-Agent Systems
2024-10-24
alphaXiv
arXiv
S2