← People
Zhijing Jin

Zhijing Jin following

University of Toronto / Vector Institute
@@ZhijingJinLinkedInGoogle ScholarWebsitePapers in the feed →

Papers · 64
  1. When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
    2026-07-06alphaXiv arXiv S2
  2. Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
    arXiv.org2026-05-08alphaXiv arXiv S2
  3. Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
    arXiv.org2026-04-17alphaXiv arXiv S2
  4. CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
    arXiv.org2026-04-16alphaXiv arXiv S2
  5. Evaluating Cooperation in LLM Social Groups through Elected Leadership
    arXiv.org2026-04-13alphaXiv arXiv S2
  6. From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
    2026-04-06alphaXiv arXiv S2
  7. Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
    arXiv.org2026-04-06alphaXiv arXiv S2
  8. When Do Language Models Endorse Limitations on Human Rights Principles?
    Conference of the European Chapter of the Association for Computational Linguistics2026-03-04alphaXiv arXiv S2
  9. Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
    arXiv.org2026-02-19alphaXiv arXiv S2
  10. GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
    arXiv.org2026-02-12alphaXiv arXiv S2
  11. IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery
    arXiv.org2026-02-08alphaXiv arXiv S2
  12. Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
    Conference of the European Chapter of the Association for Computational Linguistics2026-01-29alphaXiv arXiv S2
  13. Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
    Conference of the European Chapter of the Association for Computational Linguistics2026S2
  14. Taming Object Hallucinations with Verified Atomic Confidence Estimation
    Conference of the European Chapter of the Association for Computational Linguistics2026S2
  15. NLP for Social Good: A Survey and Outlook of Challenges, Opportunities and Responsible Deployment
    Conference of the European Chapter of the Association for Computational Linguistics2026S2
  16. Are LLMs Good Safety Agents or a Propaganda Engine?
    arXiv.org2025-11-28alphaXiv arXiv S2
  17. Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
    arXiv.org2025-11-13alphaXiv arXiv S2
  18. SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
    arXiv.org2025-10-06alphaXiv arXiv S2
  19. Test of Time: Rethinking Temporal Signal of Benchmark Contamination
    Annual Meeting of the Association for Computational Linguistics2025-08-26alphaXiv arXiv S2
  20. CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
    Conference of the European Chapter of the Association for Computational Linguistics2025-08-16alphaXiv arXiv S2
  21. Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
    arXiv.org2025-08-06alphaXiv arXiv S2
  22. When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
    Annual Meeting of the Association for Computational Linguistics2025-07-18alphaXiv arXiv S2
  23. Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
    Conference on Empirical Methods in Natural Language Processing2025-07-07alphaXiv arXiv S2
  24. Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
    arXiv.org2025-06-29alphaXiv arXiv S2
  25. Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
    Conference on Empirical Methods in Natural Language Processing2025-06-28alphaXiv arXiv S2
  26. Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
    arXiv.org2025-06-15alphaXiv arXiv S2
  27. Improving Large Language Model Safety with Contrastive Representation Learning
    Conference on Empirical Methods in Natural Language Processing2025-06-13alphaXiv arXiv S2
  28. Can Theoretical Physics Research Benefit from Language Agents?
    arXiv.org2025-06-06alphaXiv arXiv S2
  29. Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
    arXiv.org2025-05-29alphaXiv arXiv S2
  30. NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment
    2025-05-28alphaXiv arXiv S2
  31. Are Language Models Consequentialist or Deontological Moral Reasoners?
    Conference on Empirical Methods in Natural Language Processing2025-05-27alphaXiv arXiv S2
  32. When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
    arXiv.org2025-05-25alphaXiv arXiv S2
  33. Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
    2025-05-22alphaXiv arXiv S2
  34. Causality for Natural Language Processing
    arXiv.org2025-04-20alphaXiv arXiv S2
  35. Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by Everyone
    AAAI Conference on Artificial Intelligence2025-04-11S2
  36. How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
    Conference of the European Chapter of the Association for Computational Linguistics2025-03-20alphaXiv arXiv S2
  37. DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal
    Annual Meeting of the Association for Computational Linguistics2025-03-18alphaXiv arXiv S2
  38. Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
    Annual Meeting of the Association for Computational Linguistics2025-03-07alphaXiv arXiv S2
  39. Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
    arXiv.org2025-02-28alphaXiv arXiv S2
  40. Causality can systematically address the monsters under the bench(marks)
    arXiv.org2025-02-07alphaXiv arXiv S2
  41. NLP for Social Good: A Survey of Challenges, Opportunities, and Responsible Deployment
    arXiv.org2025S2
  42. Navigating Ethical Challenges in NLP: Hands-on strategies for students and researchers
    Annual Meeting of the Association for Computational Linguistics2025S2
  43. Causal Responsibility Attribution for Human-AI Collaboration
    arXiv.org2024-11-05alphaXiv arXiv S2
  44. Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
    arXiv.org2024-10-09alphaXiv arXiv S2
  45. How developments in natural language processing help us in understanding human behaviour
    Nature Human Behaviour2024-10-01S2
  46. On the Causal Nature of Sentiment Analysis
    arXiv.org2024S2
  47. Multilingual Trolley Problems for Language Models
    arXiv.org2024S2
  48. Cooperate or Collapse: Emergence of Sustainability Behaviors in a Society of LLM Agents
    arXiv.org2024S2
  49. CausalQuest: Collecting Natural Causal Questions for AI Agents
    arXiv.org2024S2
  50. NL2FOL: Translating Natural Language to First-Order Logic for Logical Fallacy Detection
    arXiv.org2024S2
  51. Moûsai: Efficient Text-to-Music Diffusion Models
    Annual Meeting of the Association for Computational Linguistics2024S2
  52. Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion
    arXiv.org2023S2
  53. ALERT: Adapt Language Models to Reasoning Tasks
    Annual Meeting of the Association for Computational Linguistics2023S2
  54. A PhD Student's Perspective on Research in NLP in the Era of Very Large Language Models
    arXiv.org2023S2
  55. Navigating the Ocean of Biases: Political Bias Attribution in Language Models via Causal Structures
    arXiv.org2023S2
  56. CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
    Neural Information Processing Systems2023S2
  57. Robotics Physically Plausible 4D Reconstruction from Monocular Videos
    2023S2
  58. Causal AI Scientist: Facilitating Causal Data Science with Large Language Models
    S2
  59. AI Poses Risks to Democratic and Social Systems
    S2
  60. Cooperate or Collapse: Emergence of Sustainability in a Society of LLM Agents
    S2
  61. Can Theoretical Physics Research Benefit from Language Agents?
    S2
  62. GOVSIM-ELECT: Elections in AI Societies
    S2
  63. When Do Language Models Endorse Limitations on Universal Human Rights Principles?
    S2
  64. Tutorial Proposal: Causality for Large Language Models
    S2