Yu Su following
Ohio State University
Papers · 39
-
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
-
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
-
Multimodal Depression Detection Through Conversational Interactions with an Emotion-Aware Social Robot: Pilot Study
JMIR Formative Research2026-04-27S2
-
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
-
Agent Learning via Early Experience
-
Watch and Learn: Learning to Use Computers from Online Videos
-
WebGuard: Building a Generalizable Guardrail for Web Agents
-
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
-
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
Annual Meeting of the Association for Computational Linguistics2025-06-05alphaXiv
arXiv
S2
-
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
-
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
-
ARM: Adaptive Reasoning Model
-
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
North American Chapter of the Association for Computational Linguistics2025-04-28alphaXiv
arXiv
S2
-
Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
Annual Meeting of the Association for Computational Linguistics2025-04-21alphaXiv
arXiv
S2
-
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
-
An Illusion of Progress? Assessing the Current State of Web Agents
-
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
-
Towards Understanding Graphical Perception in Large Multimodal Models
-
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
-
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
-
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
Annual Meeting of the Association for Computational Linguistics2025-02-17alphaXiv
arXiv
S2
-
Interpretable and Testable Vision Features via Sparse Autoencoders
-
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
-
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
-
Static Segmentation by Tracking: A Frustratingly Label-Efficient Approach to Fine-Grained Segmentation
-
Discovering Severe Adverse Reactions from Pharmacokinetic Drug-Drug Interactions through Literature Analysis and Electronic Health Record
Verification
Clinical pharmacology and therapy2024-11-25S2
-
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
-
Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
-
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
-
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
International Conference on Learning Representations2024-10-07alphaXiv
arXiv
S2
-
Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers
International Conference on Learning Representations2024-10-03alphaXiv
arXiv
S2
-
Fine-Tuning is Fine, if Calibrated
-
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Annual Meeting of the Association for Computational Linguistics2024-09-04alphaXiv
arXiv
S2
-
VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images
-
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
-
WebOlympus: An Open Platform for Web Agents on Live Websites
Conference on Empirical Methods in Natural Language Processing2024S2
-
Language Agents: Foundations, Prospects, and Risks
Conference on Empirical Methods in Natural Language Processing2024S2
-
Adaptive Chameleon or Stubborn Sloth: Unraveling the Behavior of Large Language Models in Knowledge Clashes
-
Towards a Robust and Generalizable Embodied Agent