# CS294/194-280 Advanced Large Language Model Agents | CS 294/194-280 Advanced Large Language Model Agents

https://rdi.berkeley.edu/adv-llm-agents/sp25

[Skip to the content.](https://rdi.berkeley.edu/adv-llm-agents/sp25#content)
# CS294/194-280   
Advanced Large Language Model Agents
## Spring 2025
### Prospective Students
  * **_Students interested in the course should first try enrolling in the course in CalCentral. The class number for CS194-280 is 33840. The class number for CS294-280 is 33841. Please join the waitlist if the class is full._**
  * **_We plan to expand the class size to allow more students to join. Please fill in the[petition form](https://forms.gle/sfWW8M2w1LDTnQWm9) if you are on the waitlist or can’t get added to the waitlist. You will receive an email notification around the beginning of the spring semester if you are allowed in._**
  * **_Do not email course staff or TAs. Please use[Edstem](https://edstem.org/us/join/QMmJkA) for any questions. For private matters, post a private question on Edstem and make sure it is visable to all teaching staff._**


## Course Staff  
| Instructor  | (Guest) Co-instructor  | (Guest) Co-instructor  |  
| --- | --- | --- |  
| ![](https://rdi.berkeley.edu/adv-llm-agents/assets/dawn-berkeley.jpg)  | ![](https://rdi.berkeley.edu/adv-llm-agents/assets/XinyunChen.jpg)  | ![](https://rdi.berkeley.edu/adv-llm-agents/assets/KaiyuYang.jpg)  |  
| [Dawn Song](https://people.eecs.berkeley.edu/~dawnsong/)  | Xinyun Chen  | Kaiyu Yang  |  
| Professor, UC Berkeley  | Research Scientist,   
Google DeepMind  | Research Scientist,   
Meta FAIR  |  
Teaching Staff: Alex Pan, Tara Pande, Ashwin Dara, Jason Yan
## Class Time and Location
Lecture: 4-6pm PT Monday at Anthro/Art Building 160
## Course Description
Large language model (LLM) agents have been an important frontier in AI, however, they still fall short critical skills, such as complex reasoning and planning, for solving hard problems and enabling end-to-end applications in real-world scenarios. Building on our [previous course](https://llmagents-learning.org/f24), this course dives deeper into advanced topics in LLM agents, focusing on reasoning, AI for mathematics, code generation, and program verification. We begin by introducing advanced inference and post-training techniques for building LLM agents that can search and plan. Then, we focus on two application domains: mathematics and programming. We study how LLMs can be used to prove mathematical theorems, as well as generate and reason about computer programs. Specifically, we will cover the following topics:
  * Inference-time techniques for reasoning
  * Post-training methods for reasoning
  * Search and planning
  * Agentic workflow, tool use, and functional calling
  * LLMs for code generation and verification
  * LLMs for mathematics: data curation, continual pretraining, and finetuning
  * LLM agents for theorem proving and autoformalization


## Syllabus  
| Date  | Guest Lecture   
(4:00PM-6:00PM PT)  | Supplemental Readings  |  
| --- | --- | --- |  
| Jan 27th  |  **Inference-Time Techniques for LLM Reasoning**   
Xinyun Chen, Google DeepMind   
[Recording](https://www.youtube.com/live/g0Dwtf3BH-0) [Intro](https://rdi.berkeley.edu/adv-llm-agents/slides/llm-agents-berkeley-intro-sp25.pdf) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/inference_time_techniques_lecture_sp25.pdf)  | - [Large Language Models as Optimizers](https://arxiv.org/abs/2309.03409)   
- [Large Language Models Cannot Self-Correct Reasoning Yet](https://arxiv.org/abs/2310.01798)   
- [Teaching Large Language Models to Self-Debug](https://arxiv.org/abs/2304.05128)   
_All readings are optional this week._  |  
| Feb 3rd  |  **Learning to reason with LLMs**   
Jason Weston, Meta   
[Recording](https://www.youtube.com/live/_MNlLhU33H0) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/Jason-Weston-Reasoning-Alignment-Berkeley-Talk.pdf)  | - [Direct Preference Optimization: Your Language Model is Secretly a Reward Model](https://arxiv.org/abs/2305.18290)   
- [Iterative Reasoning Preference Optimization](https://arxiv.org/abs/2404.19733)   
- [Chain-of-Verification Reduces Hallucination in Large Language Models](https://arxiv.org/abs/2309.11495)  |  
| Feb 10th  |  **On Reasoning, Memory, and Planning of Language Agents**   
Yu Su, Ohio State University   
[Recording](https://www.youtube.com/live/zvI4UN2_i-w) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/language_agents_YuSu_Berkeley.pdf)  | - [Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization](https://arxiv.org/abs/2405.15071)   
- [HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models](https://arxiv.org/abs/2405.14831)   
- [Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents](https://arxiv.org/abs/2411.06559)  |  
| Feb 17th  | _No Class - Presidents’ Day_  |   |  
| Feb 24th  |  **Open Training Recipes for Reasoning in Language Models**   
Hanna Hajishirzi, University of Washington   
[Recording](https://www.youtube.com/live/cMiu3A7YBks) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/OLMo-Tulu-Reasoning-Hanna.pdf)  | - [Tulu 3: Pushing Frontiers in Open Language Model Post-Training](https://arxiv.org/abs/2411.15124)   
- [Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback](https://arxiv.org/abs/2406.09279)   
- [OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs](https://arxiv.org/abs/2411.14199)  |  
| Mar 3rd  |  **Coding Agents and AI for Vulnerability Detection**   
Charles Sutton, Google DeepMind   
[Recording](https://www.youtube.com/live/JCk6qJtaCSU) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/Code%20Agents%20and%20AI%20for%20Vulnerability%20Detection.pdf)  | - [Interactive Tools Substantially Assist LM Agents in Finding Security Vulnerabilities](https://arxiv.org/abs/2409.16165)   
- [From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code](https://googleprojectzero.blogspot.com/2024/10/from-naptime-to-big-sleep.html)  |  
| Mar 10th  |  **Multimodal Autonomous AI Agents**   
Ruslan Salakhutdinov, CMU/Meta   
[Recording](https://www.youtube.com/live/RPINOYM12RU) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/ruslan-multimodal.pdf)  | - [Mind2Web: Towards a Generalist Agent for the Web](https://arxiv.org/abs/2306.06070)   
- [WebArena: A Realistic Web Environment for Building Autonomous Agents](https://arxiv.org/abs/2307.13854)   
- [VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks](https://jykoh.com/vwa)   
- [Tree Search for Language Model Agents](https://jykoh.com/search-agents)  |  
| Mar 17th  |  **Multimodal Agents – From Perception to Action**   
Caiming Xiong, Salesforce AI Research   
[Recording](https://www.youtube.com/live/n__Tim8K2IY) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/Multimodal_Agent_caiming.pdf)  | - [OSWORLD: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments](https://arxiv.org/pdf/2404.07972)   
- [AGUVIS: Unified Pure Vision Agents For Autonomous GUI Interaction](https://arxiv.org/pdf/2412.04454)  |  
| Mar 24th  | _No Class - Spring Recess_  |   |  
| Mar 31st  |  **AlphaProof: when reinforcement learning meets formal mathematics**   
Thomas Hubert, Google DeepMind   
10am-noon PT   
[Recording](https://www.youtube.com/live/3gaEMscOMAU) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/alphaproof.pdf)  | - [AI achieves silver-medal standard solving International Mathematical Olympiad problems](https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/)   
- [Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm](https://arxiv.org/pdf/1712.01815)   
- [The Future of Mathematics?](https://www.youtube.com/watch?v=Dp-mQ3HxgDE)   
- [Building the Mathematical Library of the Future](https://www.quantamagazine.org/building-the-mathematical-library-of-the-future-20201001/)  |  
| Apr 7th  |  **Language models for autoformalization and theorem proving**   
Kaiyu Yang, Meta FAIR   
[Recording](https://www.youtube.com/live/cLhWEyMQ4mQ) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/mathverification.pdf)  | - [LeanDojo: Theorem Proving with Retrieval-Augmented Language Models](https://arxiv.org/abs/2306.15626)   
- [Autoformalization with Large Language Models](https://arxiv.org/abs/2205.12615)   
- [Autoformalizing Euclidean Geometry](https://arxiv.org/abs/2405.17216)  |  
| Apr 14th  |  **Advanced topics in theorem proving**   
Sean Welleck, CMU   
[Recording](https://www.youtube.com/live/Gy5Nm17l9oo) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/welleck2025_berkeley_bridging.pdf)  | - [Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs](https://arxiv.org/abs/2210.12283)   
- [miniCTX: Neural Theorem Proving with Long-Contexts](https://www.arxiv.org/pdf/2408.03350)   
- [Lean-STaR: Learning to Interleave Thinking and Proving](https://arxiv.org/abs/2407.10040)   
- [ImProver: Agent-Based Automated Proof Optimization](https://arxiv.org/abs/2410.04753)  |  
| Apr 21st  |  **Abstraction and Discovery with Large Language Model Agents**   
Swarat Chaudhuri, UT Austin   
10am-noon PT   
[Recording](https://www.youtube.com/live/IHc0TEMrEdY) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/swarat.pdf)  | - [An In-Context Learning Agent for Formal Theorem-Proving](https://arxiv.org/abs/2310.04353)   
- [Symbolic Regression with a Learned Concept Library](https://arxiv.org/abs/2409.09359)  |  
| Apr 28th  |  **Towards building safe and secure agentic AI**   
Dawn Song, UC Berkeley   
[Recording](https://www.youtube.com/live/ti6yPE2VPZc) [Slides](https://rdi.berkeley.edu/adv-llm-agents/slides/dawn-agentic-ai.pdf)  | - [Privtrans: Automatically Partitioning Programs for Privilege Separation](https://dawnsong.io/papers/privtrans.pdf)   
- [DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks](https://arxiv.org/abs/2504.11358)   
- [AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases](https://arxiv.org/abs/2407.12784)   
- [Progent: Programmable Privilege Control for LLM Agents](https://arxiv.org/html/2504.11703v1)  |  
## Enrollment and Grading
**_Prerequisites:_** **Students are strongly encouraged to have had experience and basic understanding of Machine Learning and Deep Learning before taking this class, e.g., have taken courses such as CS182, CS188, and CS189.**
**_Please fill out the[petition form](https://forms.gle/sfWW8M2w1LDTnQWm9) if you are on the waitlist or can’t get added to the waitlist._**
This is a variable-unit course. All enrolled students are expected to participate in lectures in person and complete weekly reading summaries related to the course content. Students enrolling in one unit are expected to submit an article that summarizes one of the lectures. Students enrolling in more than one unit are expected to submit a lab assignment and a project instead of the article. For students enrolling in 2 units, the project should have a written report, which can be a survey in a certain area related to LLMs. For students enrolling in 3 or 4 units, projects will follow either an applications track or a research track:
  * **Applications Track:** Projects in this track focus on applied use cases of LLMs and do not necessarily need to contribute novel research. Students in this track will work in groups of 3-4. The project for 3-unit students should include an implementation (coding) component that programmatically interacts with LLMs, while 4-unit students must complete a more substantial implementation with the potential for real-world impact.
  * **Research Track:** Students in this track will conduct novel research under the supervision of postdocs and graduate students, with the goal of publishing in a workshop or conference. Research track projects must be completed in groups of 2-3, and students must apply to participate via a forthcoming Google form. The expectations for implementation and intellectual contributions will align with the project requirements for 3- and 4-unit students.


The grade breakdowns for students enrolled in different units are the following:  
|   | 1 unit  | 2 units  | 3/4 units  |  
| --- | --- | --- | --- |  
| Participation  | 40%  | 16%  | 8%  |  
| Reading Summaries  | 10%  | 4%  | 2%  |  
| Quizzes  | 10%  | 4%  | 2%  |  
| Article  | 40%  |   |   |  
| Lab  |   | 16%  | 8%  |  
| Project  |   |   |   |  
|   _Proposal_  |   | 10%  | 10%  |  
|   _Milestone_  |   | 10%  | 10%  |  
|   _Poster Presentation_  |   | 10%  | 10%  |  
|   _Presentation Recording_  |   | 10%  | 5%  |  
|   _Report_  |   | 20%  | 20%  |  
|   _Implementation_  |   |   | 25%  |  
## Lab and Project Timeline  
|   | Released  | Due  |  
| --- | --- | --- |  
| Project group formation  | 1/27  | 2/24  |  
| Project proposal  | 2/3  | 2/24  |  
| Project milestone  | 2/24  | 3/31  |  
| Lab  | 3/31  | 4/28  |  
| Project final poster presentation  | 4/28  | 5/5  |  
| Project final presentation recording  | 4/28  | 5/16  |  
| Project final report  | 4/28  | 5/16  |  
## Office Hours
  * Alex: 6-7pm on Mondays on [Zoom](https://berkeley.zoom.us/j/2012565201)

This page was generated by [GitHub Pages](https://pages.github.com).
