CS PhD · Duke University

Junlin Wang

Hi, I am Junlin, a fourth-year Computer Science PhD student at Duke University, advised by Prof. Bhuwan Dhingra, and previously also by Prof. Sam Wiseman (2022–2023). Before Duke, I worked closely with Sameer Singh on machine learning interpretability and NLP. I have also been a research intern at Google DeepMind, Together AI and AWS.

My main research goal is recursive self-improvement (RSI): AI that automates AI research. Toward that, I work on:

  • Multi-agent scaling
  • Continual learning in the real world that creates economic value
  • The science of post-training
  • Long-horizon evaluation that is robust, fast, and hard to reward-hack

News

  • Achieved best public score on ARC-AGI-3! [X] The work is now out as PRO-LONG [code] [thread].
  • I will be presenting at NeurIPS 2025!
  • Summer 2025: Research internship at Google DeepMind.
  • I will be presenting at EMNLP 2024!

Publications

  • PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

    Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra

    arXiv 2026

  • DSGym: A Holistic Framework for Evaluating and Training Data Science Agents

    Fan Nie, Junlin Wang, Harper Hua, Federico Bianchi, Yongchan Kwon, Zhenting Qi, Owen Queen, Shang Zhu, James Zou

    arXiv 2026

  • When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework

    Zhen Xu, Shang Zhu, Jue Wang, Junlin Wang, Ben Athiwaratkun, Chi Wang, James Zou, Ce Zhang

    ICLR 2026

    Paper
  • Improving Model Alignment Through Collective Intelligence of Open-Source LLMs

    Junlin Wang, Roy Xie, Shang Zhu, Jue Wang, Ben Athiwaratkun, Bhuwan Dhingra, Shuaiwen Leon Song, Ce Zhang, James Zou

    ICML 2025

  • Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods

    Junlin Wang, Shang Zhu, Jon Saad-Falcon, Ben Athiwaratkun, Qingyang Wu, Jue Wang, Shuaiwen Leon Song, Ce Zhang, Bhuwan Dhingra, James Zou

    arXiv 2025

    PaperX
  • How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

    Hongyi Cai, Junlin Wang, Xiaoyin Chen, Bhuwan Dhingra

    arXiv 2025

    PaperX
  • Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals

    Roy Xie, Junlin Wang, Paul Rosu, Chunyuan Deng, Bolun Sun, Zihao Lin, Bhuwan Dhingra

    NeurIPS 2025

  • Mixture-of-Agents Enhances Large Language Model Capabilities

    Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, James Zou

    ICLR 2025

  • Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies

    Junlin Wang, Siddhartha Jain, Dejiao Zhang, Baishakhi Ray, Varun Kumar, Ben Athiwaratkun

    EMNLP 2024

    Paper
  • ReCaLL: Membership Inference via Relative Conditional Log-Likelihoods

    Roy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang, Rong Ge, Jian Pei, Neil Zhenqiang Gong, Bhuwan Dhingra

    EMNLP 2024

  • Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications

    Junlin Wang*, Tianyi Yang*, Roy Xie, Bhuwan Dhingra

    ACL 2024 Findings

    PaperCode
  • LLM-Resistant Math Word Problem Generation via Adversarial Attacks

    Roy Xie, Chengxuan Huang, Junlin Wang, Bhuwan Dhingra

    EMNLP 2024 Findings

    PaperCode
  • NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge

    Phillip Howard*, Junlin Wang*, Vasudev Lal, Gadi Singer, Yejin Choi, Swabha Swayamdipta

    NAACL 2024 Findings

    Paper
  • Maestro: A Gamified Platform for Teaching AI Robustness

    Margarita Geleta, Jiacen Xu, Manikanta Loya, Junlin Wang, Sameer Singh, Zhou Li and Sergio Gago Masague

    EAAI 2023

    Paper
  • Gradient-based Analysis of NLP Models is Manipulable

    Junlin Wang*, Jens Tuyls*, Eric Wallace and Sameer Singh

    EMNLP 2020 Findings

    PaperBlog
  • AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models

    Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh

    Demo at EMNLP 2019Best Demo Award

    PaperBlog

Projects

Some of my for-fun projects →