← People
Zhijing Jin
following
University of Toronto / Vector Institute
@@ZhijingJin
LinkedIn
Google Scholar
Website
Papers in the feed →
Papers · 64
When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
2026-07-06
alphaXiv
arXiv
S2
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
arXiv.org
2026-05-08
alphaXiv
arXiv
S2
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
arXiv.org
2026-04-17
alphaXiv
arXiv
S2
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
arXiv.org
2026-04-16
alphaXiv
arXiv
S2
Evaluating Cooperation in LLM Social Groups through Elected Leadership
arXiv.org
2026-04-13
alphaXiv
arXiv
S2
From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
2026-04-06
alphaXiv
arXiv
S2
Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
arXiv.org
2026-04-06
alphaXiv
arXiv
S2
When Do Language Models Endorse Limitations on Human Rights Principles?
Conference of the European Chapter of the Association for Computational Linguistics
2026-03-04
alphaXiv
arXiv
S2
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
arXiv.org
2026-02-19
alphaXiv
arXiv
S2
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
arXiv.org
2026-02-12
alphaXiv
arXiv
S2
IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery
arXiv.org
2026-02-08
alphaXiv
arXiv
S2
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
Conference of the European Chapter of the Association for Computational Linguistics
2026-01-29
alphaXiv
arXiv
S2
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
Conference of the European Chapter of the Association for Computational Linguistics
2026
S2
Taming Object Hallucinations with Verified Atomic Confidence Estimation
Conference of the European Chapter of the Association for Computational Linguistics
2026
S2
NLP for Social Good: A Survey and Outlook of Challenges, Opportunities and Responsible Deployment
Conference of the European Chapter of the Association for Computational Linguistics
2026
S2
Are LLMs Good Safety Agents or a Propaganda Engine?
arXiv.org
2025-11-28
alphaXiv
arXiv
S2
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
arXiv.org
2025-11-13
alphaXiv
arXiv
S2
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
arXiv.org
2025-10-06
alphaXiv
arXiv
S2
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
Annual Meeting of the Association for Computational Linguistics
2025-08-26
alphaXiv
arXiv
S2
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
Conference of the European Chapter of the Association for Computational Linguistics
2025-08-16
alphaXiv
arXiv
S2
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
arXiv.org
2025-08-06
alphaXiv
arXiv
S2
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
Annual Meeting of the Association for Computational Linguistics
2025-07-18
alphaXiv
arXiv
S2
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
Conference on Empirical Methods in Natural Language Processing
2025-07-07
alphaXiv
arXiv
S2
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
arXiv.org
2025-06-29
alphaXiv
arXiv
S2
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
Conference on Empirical Methods in Natural Language Processing
2025-06-28
alphaXiv
arXiv
S2
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
arXiv.org
2025-06-15
alphaXiv
arXiv
S2
Improving Large Language Model Safety with Contrastive Representation Learning
Conference on Empirical Methods in Natural Language Processing
2025-06-13
alphaXiv
arXiv
S2
Can Theoretical Physics Research Benefit from Language Agents?
arXiv.org
2025-06-06
alphaXiv
arXiv
S2
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
arXiv.org
2025-05-29
alphaXiv
arXiv
S2
NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment
2025-05-28
alphaXiv
arXiv
S2
Are Language Models Consequentialist or Deontological Moral Reasoners?
Conference on Empirical Methods in Natural Language Processing
2025-05-27
alphaXiv
arXiv
S2
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
arXiv.org
2025-05-25
alphaXiv
arXiv
S2
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
2025-05-22
alphaXiv
arXiv
S2
Causality for Natural Language Processing
arXiv.org
2025-04-20
alphaXiv
arXiv
S2
Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by Everyone
AAAI Conference on Artificial Intelligence
2025-04-11
S2
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
Conference of the European Chapter of the Association for Computational Linguistics
2025-03-20
alphaXiv
arXiv
S2
DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal
Annual Meeting of the Association for Computational Linguistics
2025-03-18
alphaXiv
arXiv
S2
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
Annual Meeting of the Association for Computational Linguistics
2025-03-07
alphaXiv
arXiv
S2
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
arXiv.org
2025-02-28
alphaXiv
arXiv
S2
Causality can systematically address the monsters under the bench(marks)
arXiv.org
2025-02-07
alphaXiv
arXiv
S2
NLP for Social Good: A Survey of Challenges, Opportunities, and Responsible Deployment
arXiv.org
2025
S2
Navigating Ethical Challenges in NLP: Hands-on strategies for students and researchers
Annual Meeting of the Association for Computational Linguistics
2025
S2
Causal Responsibility Attribution for Human-AI Collaboration
arXiv.org
2024-11-05
alphaXiv
arXiv
S2
Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
arXiv.org
2024-10-09
alphaXiv
arXiv
S2
How developments in natural language processing help us in understanding human behaviour
Nature Human Behaviour
2024-10-01
S2
On the Causal Nature of Sentiment Analysis
arXiv.org
2024
S2
Multilingual Trolley Problems for Language Models
arXiv.org
2024
S2
Cooperate or Collapse: Emergence of Sustainability Behaviors in a Society of LLM Agents
arXiv.org
2024
S2
CausalQuest: Collecting Natural Causal Questions for AI Agents
arXiv.org
2024
S2
NL2FOL: Translating Natural Language to First-Order Logic for Logical Fallacy Detection
arXiv.org
2024
S2
Moûsai: Efficient Text-to-Music Diffusion Models
Annual Meeting of the Association for Computational Linguistics
2024
S2
Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion
arXiv.org
2023
S2
ALERT: Adapt Language Models to Reasoning Tasks
Annual Meeting of the Association for Computational Linguistics
2023
S2
A PhD Student's Perspective on Research in NLP in the Era of Very Large Language Models
arXiv.org
2023
S2
Navigating the Ocean of Biases: Political Bias Attribution in Language Models via Causal Structures
arXiv.org
2023
S2
CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
Neural Information Processing Systems
2023
S2
Robotics Physically Plausible 4D Reconstruction from Monocular Videos
2023
S2
Causal AI Scientist: Facilitating Causal Data Science with Large Language Models
S2
AI Poses Risks to Democratic and Social Systems
S2
Cooperate or Collapse: Emergence of Sustainability in a Society of LLM Agents
S2
Can Theoretical Physics Research Benefit from Language Agents?
S2
GOVSIM-ELECT: Elections in AI Societies
S2
When Do Language Models Endorse Limitations on Universal Human Rights Principles?
S2
Tutorial Proposal: Causality for Large Language Models
S2