← People
Chen Zhao
following
NYU Shanghai
@henryzhao4321
Papers in the feed →
Papers · 35
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
arXiv.org
2026-06-22
alphaXiv
arXiv
S2
MMGist: A Comprehensive Multimodal Benchmark for 2027
arXiv.org
2026-06-21
alphaXiv
arXiv
S2
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
arXiv.org
2026-05-29
alphaXiv
arXiv
S2
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
Annual Meeting of the Association for Computational Linguistics
2026-05-05
alphaXiv
arXiv
S2
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
arXiv.org
2026-03-10
alphaXiv
arXiv
S2
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
Annual Meeting of the Association for Computational Linguistics
2026-03-09
alphaXiv
arXiv
S2
RPDR: A Round-trip Prediction-Based Data Augmentation Framework for Long-Tail Question Answering
Conference on Empirical Methods in Natural Language Processing
2026-02-19
alphaXiv
arXiv
S2
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
arXiv.org
2026-02-05
alphaXiv
arXiv
S2
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
arXiv.org
2026-01-19
alphaXiv
arXiv
S2
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
Annual Meeting of the Association for Computational Linguistics
2026-01-19
alphaXiv
arXiv
S2
Deconstructing Multimodal Mathematical Reasoning: Towards a Unified Perception-Alignment-Reasoning Paradigm
arXiv.org
2026
S2
LimRank: Less is More for Reasoning-Intensive Information Reranking
Conference on Empirical Methods in Natural Language Processing
2025-10-27
alphaXiv
arXiv
S2
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
Conference on Empirical Methods in Natural Language Processing
2025-10-17
alphaXiv
arXiv
S2
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
arXiv.org
2025-10-10
alphaXiv
arXiv
S2
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
Conference on Empirical Methods in Natural Language Processing
2025-10-07
alphaXiv
arXiv
S2
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
arXiv.org
2025-10-07
alphaXiv
arXiv
S2
PQReCS: Accelerating Code Search with Product Quantization and Result Re-ranking Optimization
2025 5th International Conference on Artificial Intelligence, Automation and High Performance Computing (AIAHPC)
2025-09-19
S2
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
arXiv.org
2025-08-03
alphaXiv
arXiv
S2
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Neural Information Processing Systems
2025-07-01
alphaXiv
arXiv
S2
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
arXiv.org
2025-06-05
alphaXiv
arXiv
S2
Inter-Passage Verification for Multi-evidence Multi-answer QA
Annual Meeting of the Association for Computational Linguistics
2025-05-31
alphaXiv
arXiv
S2
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
Conference on Empirical Methods in Natural Language Processing
2025-05-21
alphaXiv
arXiv
S2
Physics: Benchmarking Foundation Models on University-Level Physics Problem Solving
Annual Meeting of the Association for Computational Linguistics
2025-03-26
alphaXiv
arXiv
S2
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
Conference on Empirical Methods in Natural Language Processing
2025-03-26
alphaXiv
arXiv
S2
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Computer Vision and Pattern Recognition
2025-01-21
alphaXiv
arXiv
S2
SciSketch: An Open-source Framework for Automated Schematic Diagram Generation in Scientific Papers
Conference on Empirical Methods in Natural Language Processing
2025
S2
Are Multimodal LLMs Robust Against Adversarial Perturbations? RoMMath: A Systematic Evaluation on Multimodal Math Reasoning
North American Chapter of the Association for Computational Linguistics
2025
S2
SportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question Answering
Conference on Empirical Methods in Natural Language Processing
2025
S2
MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering
Conference on Empirical Methods in Natural Language Processing
2025
S2
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents
Conference on Empirical Methods in Natural Language Processing
2024-11-08
alphaXiv
arXiv
S2
SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA
Conference on Empirical Methods in Natural Language Processing
2024-09-25
alphaXiv
arXiv
S2
KnowledgeFMath: A Knowledge-Intensive Math Reasoning Dataset in Finance Domains
Annual Meeting of the Association for Computational Linguistics
2024
S2
TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning
Annual Meeting of the Association for Computational Linguistics
2024
S2
Are MLLMs Robust Against Adversarial Perturbations? R O MM ATH : A Systematic Evaluation on Multimodal Math Reasoning
S2
P HYSICS : B ENCHMARKING F OUNDATION M ODELS FOR P ROBLEM S OLVING IN P HYSICS
S2