← People
Yu Su

Yu Su following

Ohio State University
@ysu_nlpPapers in the feed →

Papers · 39
  1. OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
    arXiv.org2026-06-28alphaXiv arXiv S2
  2. Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
    arXiv.org2026-04-27alphaXiv arXiv S2
  3. Multimodal Depression Detection Through Conversational Interactions with an Emotion-Aware Social Robot: Pilot Study
    JMIR Formative Research2026-04-27S2
  4. Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
    arXiv.org2025-10-13alphaXiv arXiv S2
  5. Agent Learning via Early Experience
    arXiv.org2025-10-09alphaXiv arXiv S2
  6. Watch and Learn: Learning to Use Computers from Online Videos
    arXiv.org2025-10-06alphaXiv arXiv S2
  7. WebGuard: Building a Generalizable Guardrail for Web Agents
    arXiv.org2025-07-18alphaXiv arXiv S2
  8. Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
    Neural Information Processing Systems2025-06-26alphaXiv arXiv S2
  9. OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
    Annual Meeting of the Association for Computational Linguistics2025-06-05alphaXiv arXiv S2
  10. BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
    Neural Information Processing Systems2025-05-29alphaXiv arXiv S2
  11. RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
    arXiv.org2025-05-28alphaXiv arXiv S2
  12. ARM: Adaptive Reasoning Model
    Neural Information Processing Systems2025-05-26alphaXiv arXiv S2
  13. MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
    North American Chapter of the Association for Computational Linguistics2025-04-28alphaXiv arXiv S2
  14. Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
    Annual Meeting of the Association for Computational Linguistics2025-04-21alphaXiv arXiv S2
  15. SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
    arXiv.org2025-04-09alphaXiv arXiv S2
  16. An Illusion of Progress? Assessing the Current State of Web Agents
    arXiv.org2025-04-02alphaXiv arXiv S2
  17. Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
    arXiv.org2025-03-31alphaXiv arXiv S2
  18. Towards Understanding Graphical Perception in Large Multimodal Models
    arXiv.org2025-03-13alphaXiv arXiv S2
  19. On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
    arXiv.org2025-02-20alphaXiv arXiv S2
  20. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
    International Conference on Machine Learning2025-02-20alphaXiv arXiv S2
  21. Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
    Annual Meeting of the Association for Computational Linguistics2025-02-17alphaXiv arXiv S2
  22. Interpretable and Testable Vision Features via Sparse Autoencoders
    2025-02-10alphaXiv arXiv S2
  23. Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
    Computer Vision and Pattern Recognition2025-01-20alphaXiv arXiv S2
  24. Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
    Computer Vision and Pattern Recognition2025-01-16alphaXiv arXiv S2
  25. Static Segmentation by Tracking: A Frustratingly Label-Efficient Approach to Fine-Grained Segmentation
    arXiv.org2025S2
  26. Discovering Severe Adverse Reactions from Pharmacokinetic Drug-Drug Interactions through Literature Analysis and Electronic Health Record Verification
    Clinical pharmacology and therapy2024-11-25S2
  27. RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
    Computer Vision and Pattern Recognition2024-11-25alphaXiv arXiv S2
  28. Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
    Trans. Mach. Learn. Res.2024-11-10alphaXiv arXiv S2
  29. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
    arXiv.org2024-10-07alphaXiv arXiv S2
  30. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
    International Conference on Learning Representations2024-10-07alphaXiv arXiv S2
  31. Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers
    International Conference on Learning Representations2024-10-03alphaXiv arXiv S2
  32. Fine-Tuning is Fine, if Calibrated
    Neural Information Processing Systems2024-09-24alphaXiv arXiv S2
  33. MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
    Annual Meeting of the Association for Computational Linguistics2024-09-04alphaXiv arXiv S2
  34. VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
    Neural Information Processing Systems2024-08-28alphaXiv arXiv S2
  35. VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
    arXiv.org2024-08-12alphaXiv arXiv S2
  36. WebOlympus: An Open Platform for Web Agents on Live Websites
    Conference on Empirical Methods in Natural Language Processing2024S2
  37. Language Agents: Foundations, Prospects, and Risks
    Conference on Empirical Methods in Natural Language Processing2024S2
  38. Adaptive Chameleon or Stubborn Sloth: Unraveling the Behavior of Large Language Models in Knowledge Clashes
    arXiv.org2023S2
  39. Towards a Robust and Generalizable Embodied Agent
    2023S2