LM4Plan @ ICML 2026
ICML’26 Workshop,
Seoul, South Korea,
Date: July 11, 2026,
Room Main Program: Grand Ballroom 101-102,
Poster Session: Hall A, poster boards 1104–1115 and 1200–1201
Invited Speakers
We are excited to announce our invited speakers for LM4Plan @ ICML 2026. View full speaker details.
- Samy Bengio - Apple
- Subbarao Kambhampati - Arizona State University
- Yarin Gal - University of Oxford
- Elias Bareinboim - Columbia University
- Noam Brown - OpenAI
- Nathan Sturtevant - University of Alberta
Overview
Language Models (LMs) are a disruptive force, changing how research was done in many subareas of AI. Planning is one of the last bastions that remain standing. The focus of this workshop is on the questions in the intersection of these areas. Some of the specific areas we would like to gain a better understanding in include: what LMs can contribute to planning, how LMs can/should be used, what are the pitfalls of using LMs, what are the guarantees that can be obtained.
Schedule
| Time | Event | Details |
|---|---|---|
| 08:00–08:05 | Opening remarks | Katharina Stein |
| 08:05–08:50 | Invited talk | Samy Bengio: Reasoning with LLMs: Challenges and Opportunities |
| 08:50–09:35 | Invited talk | Nathan Sturtevant: LLM Planning Success |
| 09:35–10:00 | Coffee break | |
| 10:00–10:45 | Invited talk | Subbarao Kambhampati: On the Role of Verifiers and Thinking Traces in Reasoning Models |
| 10:45–11:30 | Invited talk | Elias Bareinboim: Towards Causal Artificial Intelligence |
| 11:30–11:40 | buffer | |
| 11:40–12:00 | Oral presentations | LM-Landmarks: Language Model Guided Landmark Generation for Classical Planning with Formal Soundness Guarantees |
| AVATAR-AGENT: A Multi-Agent LLM Planning System for 3D Avatar | ||
| Generation | ||
| 12:00–13:00 | Lunch | |
| 13:00–13:45 | Invited talk | Yarin Gal: Model Collapse and its Implications on Planning with LLMs |
| 13:45–14:30 | Invited talk (online) | Noam Brown: Implications of Large-Scale Test-Time Compute |
| 14:30–14:40 | buffer | |
| 14:40–15:20 | Oral presentations | HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM Agents |
| ICPRL: Acquiring Physical Intuition from Interactive Control | ||
| Reward Prediction with Factorized World States | ||
| OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis | ||
| 15:20–15:30 | Concluding remarks | |
| 15:30–16:00 | Coffee break | |
| 16:00–17:00 | Poster session | Hall A, boards 1104–1115 and 1200–1201 |
Invited Talks
Elias Bareinboim:
Title: Towards Causal Artificial Intelligence
Abstract:
While a significant portion of AI scientists and engineers
believe we are on the verge of achieving highly general forms of AI, I
offer a critical appraisal of this view through a causal lens. In
particular, building on foundational developments in the field, I will
present my perspective on the relationship between intelligence and
causality, and the central role of the latter in building intelligent
systems and advancing credible data science.
I frame this discussion in terms of five core capabilities that we should expect from an intelligent AI system:
- Performing causal reasoning and articulating explanations;
- Making precise, surgical, and sample-efficient decisions;
- Generalizing across changing conditions and environments;
- Generating and simulating in a causally consistent manner; and
- Learning causal structures and variables.
In this talk, I will elaborate on this perspective and share current progress toward building causally intelligent AI systems. A more detailed discussion of this thesis is provided in my forthcoming textbook, a draft of which is available here: https://causalai-book.net/.
Samy Bengio
Title: Reasoning with LLMs: Challenges and Opportunities
Abstract: In this presentation, I will go over a few recent topics towards understanding the limits of reasoning capabilities of large language models. I will start with a discussion on what is reasoning, provide a few negative results showing limits regarding math problems and logic puzzles, and then follow up with a few more constructive approaches that can improve reasoning capabilities of LLMs.
Nathan Sturtevant
Title: LLM Planning Success
Abstract: Recent literature has shown that LMs cannot reliably plan, especially in comparison to classical planners and planning languages such as PDDL. At the same time, other work has shown that LMs are Turing complete, functionally equivalent to the computers on which they run. In this talk we present our study of this gap. We show how we are able to train LMs and achieve 99% or higher success rates across a range of planning problems such as Towers of Hanoi and Blocksworld variants on which other approaches have failed to scale. Our analysis shows where LMs are able to generalize and some cases where they are not.
Noam Brown
Title: Implications of Large-Scale Test-Time Compute
Abstract: As LLMs become more capable, their performance increasingly depends on the amount of test-time compute used. This talk argues that single-number benchmarks obscure both capability and safety-relevant trends, especially as longer chains of thought, scaffolds, and multi-agent methods push performance higher with more inference. I will close with proposals for evaluating models using performance-vs-compute curves and updating benchmarks and preparedness frameworks to account for realistic high-compute use.
Yarin Gal
Title: Model Collapse and its Implications on Planning with LLMs
Abstract: TBD
Subbarao Kambhampati
Title: On the Role of Verifiers and Thinking Traces in Reasoning Models
Abstract:
Most of the recent successes of LLMs came from the application of Reasoning Models. I will provide a perspective on reasoning models in terms of verifier-based test-time scaling methods taken to post-training. From this perspective, post-training can be viewed as laboriously compiling the verifier signal into the model weights via guessed solutions to synthetic problems. Unlike standard LLMs, reasoning models also emit the so-called “thinking traces” on the way to guessing the solution. The literature has ascribed several properties to these traces–including that they provide the end user a window into the LLM’s “thinking”, and that the length of the traces is proportional to the complexity of the problem at hand. I will share results from our recent research that call these claims into question, and provide an alternate view of the role of intermediate tokens in the reasoning models.
Important information for authors of accepted papers
See here for important information about presentations and camera-ready paper versions.
Accepted Papers
- VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
Ninghan Zhong ⋅ Ahmet E Tanriverdi ⋅ Kaan Kale ⋅ Sriram Vishwanath - AVATAR-AGENT: A Multi-Agent LLM Planning System for 3D Avatar Generation
Jason Ding ⋅ Rohan Gangaraju ⋅ Krishna C Garikipati ⋅ Foad Dabiri ⋅ Chang Xu - CAMEO: Cooperative Agentic Multi-objective Evolution of Heuristics
Tien Dat Vu ⋅ Hung Phan ⋅ Yifan Yang ⋅ Huynh Thi Thanh Binh - When Plans Collide: Joint Planning in LLM Negotiation Dyads
Yiheng Yao ⋅ Chelsea Zou ⋅ Robert Hawkins - BaRA: BFS-and-Reflection Web Data Collection Agent
Soojeong Lee ⋅ Joseph Lee ⋅ Yongseong Cho ⋅ Sunjae Kim ⋅ Youngwoo Moon ⋅ Kyungwoo Song - QueryWeaver: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation
Aishwarya Chakravarthy ⋅ Vidhi Kulkarni ⋅ Polo Chau - ICPRL: Acquiring Physical Intuition from Interactive Control
Xinrun Xu ⋅ Pi Bu ⋅ Ye Wang ⋅ Börje F. Karlsson ⋅ Ziming Wang ⋅ Tengtao Song ⋅ Qi Zhu ⋅ Jun Song ⋅ Shuo Zhang ⋅ Zhingming Ding ⋅ Bo Zheng - OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis
Naaisha Agarwal ⋅ Yihan Wu ⋅ Yichang Jian ⋅ Yikuan Hu ⋅ Nishad Mansoor ⋅ Mohan Li ⋅ Yifei Peng ⋅ Wang-Zhou Dai ⋅ Yao-Xiang Ding ⋅ Emanuele Sansone - End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent
Amin Tabrizian ⋅ Arsyi Aziz ⋅ Aarifah Ullah⋅ Mahyar Ghazanfari ⋅ Pouria Razzaghi ⋅ Peng Wei - LaGO: Latent Action Guidance for Online Reinforcement Learning
Kuanyen Liu ⋅ Renjyun Huang ⋅ Ti-Rong Wu - When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning
Nathan Gavenski ⋅ Juarez Monteiro ⋅ Francisco Galuppo Azevedo ⋅ Adriano Veloso ⋅ Odinaldo Rodrigues - Specialized LM Agents with Simulation-Verified Search for Service Workforce Planning
Vivek Singh ⋅ Santosh Pai ⋅ Sarith Mohan ⋅ Chetan L Srinidhi ⋅ Azra Aziz ⋅ Sean Cohen ⋅ Neil Biehn ⋅ Maik Kuehnhoff - IDP-MCTS Empowering Small Language Models for Automated MILP Modeling and Code Generation in Flexible Job Shop Scheduling
Mingming Peng ⋅ Jin Huang ⋅ Qihao Liu ⋅ Liang GAO ⋅ Xinyu Li - Theory of Mind Beyond Conversational Persuasion: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Ben Slater ⋅ Matteo G Mecattaf ⋅ Lucy G Cheke ⋅ John Burden ⋅ Winnie Street - SafeRun: Enabling Determinism in LLM Planning for Running
Meilin Chen ⋅ Zepeng Zhai ⋅ Jiaxuan Zhao ⋅ Yuan Lu - LM-Landmarks: Language Model Guided Landmark Generation for Classical Planning with Formal Soundness Guarantees
Kaustubh Bukkapatnam ⋅ Siddharth Karuturi - Hierarchical Chain-of-Thought: Enhancing LLM Reasoning Performance and Efficiency
Xingshuai Huang ⋅ Derek Li ⋅ Bahareh Nikpour ⋅ Parsa Omidi - The Orchestrator Bottleneck: Formal Coordination Strategies for Cost-Optimal Multi-Agent Enterprise Workflows
Rudrendu Kumar Paul ⋅ Sourav Nandy - Multi-Pass LLM Compilation for HTN Domain Authoring
Eliott Jacopin ⋅ Éric Jacopin ⋅ Koichi Takahashi - LLM-Guided Transportation Hub Capacity Planning with Textual Business Inputs
Xiaoyue Liu ⋅ Zheng Dong - Generating Robust Portfolios of Optimization Models using Large Language Models
Eleni Straitouri ⋅ Cheol Kim ⋅ Milind Tambe - HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM Agents
Jiangweizhi Peng ⋅ Yuanxin Liu ⋅ Ruida Zhou ⋅ Charles Fleming ⋅ Zhaoran Wang ⋅ Alfredo Garcia ⋅ Mingyi Hong - Calibrate Once, Choose the Beam: A Predictive Regime Test for Same-LM Search Guidance and Pruning
Jiahui Qu ⋅ Yifang Qin ⋅ BOYANG ZHENG ⋅ Ziyi Zhou - UniDesigner: Language Models as Unified Planners for Agentic Design
Zhouqiang Jiang ⋅ Bowen Wang ⋅ Shuqiong Wu ⋅ Yuta Nakashima - RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning
Tuan-Kiet Nguyen-Viet ⋅ Pham Bui Dinh ⋅ Chính Dương ⋅ Tung Dao ⋅ Cong D Tran ⋅ Huynh Thi Thanh Binh - ReTreVal: Reasoning Tree with Validation and Cross-Problem Memory for Large Language Models
Abhishek HS ⋅ Pavan C Shekar ⋅ Aswanth Krishnan ⋅ Arpit Jain - Reward Prediction with Factorized World States
Yijun Shen ⋅ Delong Chen ⋅ Xianming Hu ⋅ Jiaming Mi ⋅ Hongbo Zhao ⋅ Kai Zhang ⋅ Pascale Fung - AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
Rudrendu Kumar Paul ⋅ Sourav Nandy
Submission Instruction
The submission deadline has passed. For the previous instruction for submitting see here.
Program Committee
To be announced.
Organizing Committee
- Michael Katz, IBM
- Augusto B. Corrêa, University of Oxford
- Nir Lipovetzky, University of Melbourne
- Sarath Sreedharan, Colorado State University
- Katharina Stein, Saarland University
- Luckeciano C. Melo, University of Oxford
- Elliot Gestrin, Linköping University
Please send your inquiries to llmforplanning@gmail.com