LM4Plan @ ICAPS 2026
Date: June 29, 2026
Update: Invited talk slides and recording are now available!
Overview
Language Models (LMs) are a disruptive force, changing how research was done in many subareas of AI. Planning is one of the last bastions that remain standing. The focus of this workshop is on the questions in the intersection of these areas. Some of the specific areas we would like to gain a better understanding in include: what LMs can contribute to planning, how LMs can/should be used, what are the pitfalls of using LMs, what are the guarantees that can be obtained.
Topics of Interest
We invite paper submissions on the following (not exhaustive) list of topics:
- Planning directly with pre-trained or fine-tuned LMs.
- Planning for LMs.
- LMs for (partial) model elicitation.
- LMs for generating structured planning problem descriptions.
- LMs for search guidance or search pruning.
- LMs for validation and verification of plans, policies, or models.
- LMs for generalization in planning and generalized planning.
- Using LMs as a proxy for user preferences.
- Using LMs to develop interfaces for planning-based systems or planning-related problems.
- Other applications of LMs in planning.
We particularly encourage submissions leveraging small and open-weight language models, especially those advancing efficient and reliable methods with rigorous, reproducible evaluations and accessible research artifacts.
Important Dates
Paper submission deadline: May 1st, 2026, AoE May 8th, 2026, AoE
Paper acceptance notification: June 2nd, 2026, AoE
ICAPS will be in-person this year. Authors of accepted workshop papers are expected to register for the workshop, physically attend the conference and present in person. All accepted papers will have an oral presentation.
Submission Details
We solicit workshop paper submissions relevant to the above call. Paper submissions should be made through OpenReview.
Format
All submissions must be a single PDF file and follow one of the formats below:
Long papers – up to 8 pages + unlimited references / appendices Short papers – up to 4 pages + unlimited references / appendices
Style
Please format submissions in AAAI style (see instructions in the Author Kit ).
Dual-submission and non-archival policy
Authors submitting papers rejected from other conferences, please ensure you do your utmost to address the comments given by the reviewers. Please do not submit papers that are already accepted for the main ICAPS conference to the workshop. The workshop is a non-archival venue and will not have official proceedings. Workshop submissions can be subsequently or concurrently submitted to other venues.
Reviewing policy
Double-blind reviewing
All submissions must be anonymized and may not contain any identifying information that may violate the double-blind reviewing policy. Submissions and reviews will not be public. Only accepted papers will be made public.
Reciprocal reviewing
Depending on the number of submissions, we may adopt a reciprocal reviewing process. For each submission, one reciprocal reviewer needs to be nominated who agrees to serve as a reviewer if reciprocal reviewing is implemented. Each nominated reviewer must have at least one relevant publication at a top venue and cannot be nominated as a reviewer for more than one submission.
Invited Talk
LLMs don’t even need to plan — they can build planners

Jendrik Seipp, Linköping University
📄 Slides · ▶️ Watch the talk
Abstract
For years, LLMs couldn’t reliably solve even the smallest planning tasks. That has changed: we recently showed that the latest frontier models now beat even the strongest classical planners on several benchmark domains. Using an LLM as the planner is still rarely the best choice. Having it build planner components instead is faster, cheaper, and far less energy-hungry. I show how to do this while keeping the guarantees that make classical planning worth using, like optimality and bounded runtime. And it doesn’t stop at planners. Agents can now carry out research largely on their own, and I’ll close with some thoughts on what that means for how we work as researchers.
Speaker Bio
Jendrik Seipp is a Senior Associate Professor in Artificial Intelligence at Linköping University, Sweden, where he directs the Machine Reasoning Lab within the AIICS division. His research focuses on AI planning and its connections to machine learning. He earned his MSc in computer science from the University of Freiburg, Germany (2012) and his PhD from the University of Basel, Switzerland (2018), where he then worked as a postdoctoral researcher until 2020. He joined Linköping University as an assistant professor in 2021 and was promoted to senior associate professor in 2024. His work is supported by WASP, a Swedish Research Council Starting Grant, a Wallenberg Academy Fellowship, and an SSF Future Research Leaders grant.
Schedule
Talk length is shown in the last column in minutes. Paper talks are 10 min presentation + 5 min Q&A.
Morning
Session 1 - 9:00-10:20 Chair: Elliot Gestrin
| Time | Paper | Authors | Len |
|---|---|---|---|
| 9:00 | Opening remarks | Michael Katz | 5 |
| 9:05 | Semantic Partial Grounding via LLMs | Giuseppe Canonaco, Alberto Pozanco, Daniel Borrajo | 15 |
| 9:20 | Benchmarking LLM Pipelines for Natural Language to Automated Planning Models | Marcus Tantakoun, Christian Muise, Xiaodan Zhu | 15 |
| 9:35 | Grounded Evaluation and Repair for NL-to-PDDL Problem Generation | Joana Rosa, Bruno Martins, L. Miguel Silveira, Pedro Ricardo Leitão dos Santos | 15 |
| 9:50 | Towards LLM-Driven Synthesis of Narrative Planning Models | Allix Fletcher, Christian Muise | 15 |
| 10:05 | A Natural Language Copilot for Interactive Plan Space Exploration | Paul Horn, Daniel Gnad | 15 |
Coffee break - 10:30-10:50
Session 2 - 10:50-12:20 Chair: Nir Lipovetzky
| Time | Paper | Authors | Len |
|---|---|---|---|
| 10:50 | FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model Evaluation | Vishal Pallagani, Nitin Gupta, John A. Aydin, Biplav Srivastava | 15 |
| 11:05 | ALPSBench: Can Large Language Models Reason Their Way Through Planning Formalisms? | Marcus Tantakoun, Christian Muise, Xiaodan Zhu | 15 |
| 11:20 | On the Ability of Transformers to Verify Plans | Yash Sarrof, Yupei Du, Katharina Stein, Alexander Koller, Sylvie Thiebaux, Michael Hahn | 15 |
| 11:35 | Toward a General Framework for Evaluating Per-Domain Generalization Using LLMs, Theorem Provers, and Statistical Model Checking | Nicola J. Müller, Naya Rudolph, Ayal Taitler, Timo P. Gros | 15 |
| 11:50 | Integrating the Unified Planning Framework via MCP with Large Language Models for Reliable Automated Planning | João Areias Saraiva, Thomas Kirste | 15 |
| 12:05 | Learning HTNs from Visual Demonstration with Vision-Language Models: Preliminary Results | Ngoc La, Karthik Mahadevan, Pulkit Verma, Julie Shah | 15 |
Lunch break - 12:30-14:00
Afternoon
Session 3 - 14:00-15:15 Chair: Katharina Stein
| Time | Paper | Authors | Len |
|---|---|---|---|
| 14:00 | LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning | Elliot Gestrin, Jendrik Seipp | 15 |
| 14:15 | Personalized Medication Planning via Direct Domain Modeling and LLM-Generated Heuristics | Yonatan Vernik, David Izhaki, Alexander Tuisov, Hana Weitman, Alexander Shleyfman, Gal Kaminka | 15 |
| 14:30 | Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents | Shirin Sohrabi, Haritha Ananthakrishnan, Harsha Kokel, Kavitha Srinivas, Michael Katz | 15 |
| 14:45 | The Curious Case of Planning for Unreliable Agents: Challenges and Opportunities in Orchestrating Generative AI Agents | Roya Daneshi, Sunandita Patra, Kshama Dwarakanath, Sriram Gopalakrishnan, Daniel Borrajo, Sarath Sreedharan | 15 |
| 15:00 | Think Hierarchically, Act Optimally: Decoupled Hierarchical Planning and Execution for LLM Agents | Vikas Kumar, Jyotsana Khatri, Shirish Karande | 15 |
Coffee break - 15:30-15:50
Session 4 - 15:50-17:30 Chair: Augusto B. Corrêa
| Time | Item | Speaker | Len |
|---|---|---|---|
| 15:50 | Invited talk: LLMs don’t even need to plan — they can build planners (video) | Jendrik Seipp | 50 |
| 16:40 | Panel discussion and closing remarks | Moderator: Christian Muise, Panelists: Jendrik Seipp, Katharina Stein, Nir Lipovetzky | 50 |
Accepted Papers
- Semantic Partial Grounding via LLMs
Giuseppe Canonaco, Alberto Pozanco, Daniel Borrajo - FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model Evaluation
Vishal Pallagani, Nitin Gupta, John A. Aydin, Biplav Srivastava - ALPSBench: Can Large Language Models Reason Their Way Through Planning Formalisms?
Marcus Tantakoun, Christian Muise, Xiaodan Zhu - Benchmarking LLM Pipelines for Natural Language to Automated Planning Models
Marcus Tantakoun, Christian Muise, Xiaodan Zhu - Learning HTNs from Visual Demonstration with Vision-Language Models: Preliminary Results
Ngoc La, Karthik Mahadevan, Pulkit Verma, Julie Shah - Integrating the Unified Planning Framework via MCP with Large Language Models for Reliable Automated Planning
João Areias Saraiva, Thomas Kirste - LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning
Elliot Gestrin, Jendrik Seipp - On the Ability of Transformers to Verify Plans
Yash Sarrof, Yupei Du, Katharina Stein, Alexander Koller, Sylvie Thiebaux, Michael Hahn - Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents
Shirin Sohrabi, Haritha Ananthakrishnan, Harsha Kokel, Kavitha Srinivas, Michael Katz - Toward a General Framework for Evaluating Per-Domain Generalization Using LLMs, Theorem Provers, and Statistical Model Checking
Nicola J. Müller, Naya Rudolph, Ayal Taitler, Timo P. Gros - The Curious Case of Planning for Unreliable Agents: Challenges and Opportunities in Orchestrating Generative AI Agents
Roya Daneshi, Sunandita Patra, Kshama Dwarakanath, Sriram Gopalakrishnan, Daniel Borrajo, Sarath Sreedharan - Personalized Medication Planning via Direct Domain Modeling and LLM-Generated Heuristics
Yonatan Vernik, David Izhaki, Alexander Tuisov, Hana Weitman, Alexander Shleyfman, Gal Kaminka - Towards LLM-Driven Synthesis of Narrative Planning Models
Allix Fletcher, Christian Muise - Think Hierarchically, Act Optimally: Decoupled Hierarchical Planning and Execution for LLM Agents
Vikas Kumar, Jyotsana Khatri, Shirish Karande - A Natural Language Copilot for Interactive Plan Space Exploration
Paul Horn, Daniel Gnad - Grounded Evaluation and Repair for NL-to-PDDL Problem Generation
Joana Rosa, Bruno Martins, L. Miguel Silveira, Pedro Ricardo Leitão dos Santos
Program Committee
To be announced.
Organizing Committee
- Augusto B. Corrêa, University of Oxford
- Elliot Gestrin, Linköping University
- Sarath Sreedharan, Colorado State University
- Michael Katz, IBM
- Nir Lipovetzky, University of Melbourne
- Katharina Stein, Saarland University
- Luckeciano C. Melo, University of Oxford
Please send your inquiries to llmforplanning@gmail.com