# COS 484: Natural Language Processing

https://nlp.cs.princeton.edu/cos484/

![](https://princeton-nlp.github.io/cos484/photos/princeton-logo.png)
COS 484: Natural Language Processing
Spring 2026
[Information](https://princeton-nlp.github.io/cos484/#info) [Schedule](https://princeton-nlp.github.io/cos484/#schedule) [Coursework](https://princeton-nlp.github.io/cos484/#coursework) [FAQ](https://princeton-nlp.github.io/cos484/#faq)
  

Links
  * [Canvas web site](https://princeton.instructure.com/courses/17824)
  * [Ed](https://edstem.org/us/courses/73522/discussion): please use this for all course-related questions. You can make use of Private posts for personal matters.
  * Some useful documents:  [Working with Assignment Materials](https://docs.google.com/document/d/114CQ70qk8E--90qM8kxpTeegSATmDLIp7z_GVqJAGrk/edit?usp=sharing) [Working with LaTex](https://docs.google.com/document/d/13PIBndWXRrEzZeE70zfMooNNh6bt81QspaFz2PNdDIg/edit?usp=sharing) [Working with Google Colab](https://docs.google.com/document/d/1LlnXoOblXwW3YX-0yG_5seTXJsb3kRdMMRYqs8Qqum4/edit?usp=sharing)


What is this course about?
Recent advances have ushered in exciting developments in natural language processing (NLP), resulting in systems that can translate text, answer questions and even hold spoken conversations with us. This course will introduce students to the basics of NLP, covering standard frameworks for dealing with natural language as well as algorithms and techniques to solve various NLP problems, including recent deep learning approaches. Topics covered include language modeling, representation learning, text classification, sequence tagging, machine translation, Transformers, and others.   
  

Information
Course staff:
  * Instructors: [Tri Dao](https://tridao.me/) [Karthik Narasimhan](https://www.cs.princeton.edu/~karthikn/), 
  * TAs: [Max Gupta](https://maxdgu.github.io/), [Lucy He](https://lumos23.github.io/), [Sijia Liu](https://sijial430.github.io/), Ambri Ma, Keerthana Nallamotu, [Howard Yen](https://howard-yen.github.io/)


Time/location:
(All times are in EST) 
  * Lectures: Tuesdays/Thursdays, 10:40am-12:00pm, Frist 302
  * Precepts: Fridays 12-1pm Friend 004


Office hours for instructors and TAs are listed in the Google Calendar below.
Grading
  * **Assignments** (40%): There will be four assignments with both written and programming parts. Each homework is centered around an application and will also deepen your understanding of the theoretical concepts. 
    * Assignment 1: Language models, text classification, neural networks (10%)
    * Assignment 2: Sequence modeling, recurrent neural networks (10%)
    * Assignment 3: Transformers, attention (10%)
    * Assignment 4: Systems for LLMs, agents (10%)
  * **Midterm exam** (25%): The midterm (in person exam) will test your knowledge and problem-solving skills. 
  * **Final project** (35%): The final project offers you a chance to apply your newly acquired skills towards an in-depth application. You are required to turn in a project proposal (due on **March 20th**) and complete a paper written in the style of a conference (e.g., ACL) submission (due on **May 5th**). There will be also project presentations at the end of the semester. 
  * **Extra credit** (5%): For participation in class and Ed discussion. Limited to overall score of max 100%. 


Prerequisites:
  * Required: [COS 324](https://www.cs.princeton.edu/courses/archive/fall21/cos324/), knowledge of probability, linear algebra, multivariate calculus.
  * Proficiency in Python: programming assignments and projects will require use of Python, Numpy and PyTorch.


Reading:
There is no required textbook for this class, and you should be able to learn everything from the lectures and assignments. However, if you would like to pursue more advanced topics or get another perspective on the same material, here are some books (all of them can be read free online): 
  * Dan Jurafsky and James H. Martin. [Speech and Language Processing (3rd ed. draft).](https://web.stanford.edu/~jurafsky/slp3/)
  * Jacob Eisenstein. [Natural Language Processing ](https://github.com/jacobeisenstein/gt-nlp-class/blob/master/notes/eisenstein-nlp-notes.pdf)
  * Christopher Manning and Hinrich Schütze. [Foundations of Statistical Natural Language Processing](https://nlp.stanford.edu/fsnlp/).


Acknowledgements:
We thank [Thinking Machines Lab](https://thinkingmachines.ai/news/tinker-research-and-teaching-grants/) for providing Tinker credits for students in this course, and [GPU Mode](https://www.gpumode.com/home) and [Modal](https://modal.com/) for hosting our kernel competition leaderboard. 
Schedule
Lecture schedule is tentative and subject to change. All assignments are due at **11:59pm EST** on Tuesdays. Check the schedule for exact dates   
| Week  | Date  | Topics  | Readings  | Assignments  |  
| --- | --- | --- | --- | --- |  
| 1  | Tue (1/27)  | [Introduction + Language Models](https://princeton-nlp.github.io/cos484/lectures/L1_sp26.pdf)  |  1. [Advances in natural language processing](https://princeton-nlp.github.io/cos484/readings/advances_in_nlp.pdf)   
2. [Human Language Understanding & Reasoning](https://direct.mit.edu/daed/article/151/2/127/110621/Human-Language-Understanding-amp-Reasoning)   
3. [J & M 3.1-3.5](https://web.stanford.edu/~jurafsky/slp3/3.pdf)  |   |  
|   | Thu (1/29)  | [Text classification](https://princeton-nlp.github.io/cos484/lectures/L2_sp26.pdf)  |  [Logistic regression: J & M 4.1-4.6](https://web.stanford.edu/~jurafsky/slp3/4.pdf)   
 |   |  
|   | Fri (1/30)  | [Precept 1](https://princeton-nlp.github.io/cos484/precepts/precept1_s26.pdf)  |   |   |  
| 2  | Tue (2/3)  | [Word Embeddings](https://princeton-nlp.github.io/cos484/lectures/L3_sp26.pdf)  |  1. [Embeddings: J & M 5.1-5.8](https://web.stanford.edu/~jurafsky/slp3/5.pdf)  
2. [Efficient Estimation of Word Representations in Vector Space](https://arxiv.org/pdf/1301.3781.pdf) (word2vec)  
3. [Distributed representations of words and phrases and their compositionality](https://proceedings.neurips.cc/paper/2013/file/9aa42b31882ec039965f3c4923ce901b-Paper.pdf) (negative sampling)   | [A1 out](https://princeton-nlp.github.io/cos484/assignments/a1_s26.pdf)  |  
|   | Thu (2/5)  | [Neural Networks for NLP](https://princeton-nlp.github.io/cos484/lectures/L4_sp26.pdf)  |  [J & M 6.2-6.4, 6.6, 6.8, 6.10-6.12](https://web.stanford.edu/~jurafsky/slp3/6.pdf)  
 |   |  
|   | Fri (2/6)  | [Precept 2](https://princeton-nlp.github.io/cos484/precepts/precept2_s26.pdf)  |   |   |  
| 3  | Tue (2/10)  | [Sequence Models](https://princeton-nlp.github.io/cos484/lectures/L5_sp26.pdf)  |  1. [J&M 17.1-17.4](https://web.stanford.edu/~jurafsky/slp3/17.pdf)  
2. [Michael Collin's notes on HMMs](https://princeton-nlp.github.io/cos484/readings/hmms-spring2013.pdf)  
 |   |  
|   | Thu (2/12)  | [Recurrent neural networks](https://princeton-nlp.github.io/cos484/lectures/L6_sp26.pdf)  |  1. [J&M 13.1-13.4](https://web.stanford.edu/~jurafsky/slp3/13.pdf)  
2. [The Unreasonable Effectiveness of Recurrent Neural Networks](http://karpathy.github.io/2015/05/21/rnn-effectiveness/)  
 |   |  
|   | Fri (2/13)  | [Precept 3](https://princeton-nlp.github.io/cos484/precepts/precept3_s26.pdf)  |   |  
| 4  | Tue (2/17)  | [Seq2Seq models + Attention](https://princeton-nlp.github.io/cos484/lectures/L7_sp26.pdf)  |  1. [Sequence to Sequence Learning with Neural Networks](https://arxiv.org/pdf/1409.3215.pdf)  
2. [Neural Machine Translation by Jointly Learning to Align and Translate](https://arxiv.org/pdf/1409.0473.pdf)  
3. [Effective Approaches to Attention-based Neural Machine Translation](https://arxiv.org/pdf/1508.04025.pdf)  
4. [Blog post: Visualizing A Neural Machine Translation Model](https://jalammar.github.io/visualizing-neural-machine-translation-mechanics-of-seq2seq-models-with-attention/)  | A1 due, [A2 out](https://princeton-nlp.github.io/cos484/assignments/a2_s26.pdf)  |  
|   | Thu (2/19)  | [Transformers 1](https://princeton-nlp.github.io/cos484/lectures/L8_sp26.pdf)  |  1. [J&M 8.1-8.4](https://web.stanford.edu/~jurafsky/slp3/8.pdf)  
2. [Attention Is All You Need](https://arxiv.org/pdf/1706.03762.pdf)  
3. [The Annotated Transformer](https://nlp.seas.harvard.edu/2018/04/03/attention.html)  
4. [The Illustrated Transformer](http://jalammar.github.io/illustrated-transformer/)  |   |  
|   | Fri (2/20)  | [Precept 4](https://princeton-nlp.github.io/cos484/precepts/precept4_s26.pdf)  |   |   |  
| 5  | Tue (2/24)  |  [Transformers 2](https://princeton-nlp.github.io/cos484/lectures/L9_sp26.pdf)[ (video)](https://drive.google.com/file/d/1ur8M4zTFcb3z7k_Rc_qUcbqC_47TDHKM/view?usp=sharing)  |  1. [Efficient Transformers: A Survey](https://arxiv.org/pdf/2009.06732.pdf)  
2. [Vision Transformer](https://arxiv.org/pdf/2010.11929.pdf)  |   |  
|   | Thu (2/26)  | [Contextualized representations + Pre-Training](https://princeton-nlp.github.io/cos484/lectures/L10_sp26.pdf)  |  1. [Deep contextualized word representations](https://arxiv.org/pdf/1802.05365.pdf) (ELMo)  
2. [Improving Language Understanding by Generative Pre-Training](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf) (GPT)  
3. [BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding](https://arxiv.org/pdf/1810.04805.pdf)  
4. [The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning)](http://jalammar.github.io/illustrated-bert/)  |   |  
|   | Fri (2/27)  | [Precept 5](https://princeton-nlp.github.io/cos484/precepts/precept5_s26.pdf)  |   |   |  
| 6  | Tue (3/3)  | [Midterm Review](https://princeton-nlp.github.io/cos484/precepts/26Spring_COS484_MidtermReview.pdf)  |   | A2 due  |  
|   | Thu (3/5)  | Midterm (in person)   |   |   |  
|   | Fri (3/6)  | No precept  |   |   |  
| 7  | Tue (3/10)  | Spring break (no class)  |   | [A3 out](https://princeton-nlp.github.io/cos484/assignments/a3_s26.pdf)  |  
|   | Thu (3/12)  | Spring break (no class)  |   |   |  
| 8  | Tue (3/17)  | [Large Language Models (LLMs)](https://princeton-nlp.github.io/cos484/lectures/L11_sp26.pdf)  |  1. [Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer](https://arxiv.org/pdf/1910.10683.pdf) (T5)  
2. [Language Models are Few-Shot Learners](https://arxiv.org/pdf/2005.14165.pdf) (GPT-3)  
3. [GPT-4 Technical Report](https://arxiv.org/pdf/2303.08774.pdf) (GPT-4)   
4. [Chain-of-Thought Prompting Elicits Reasoning in Large Language Models](https://arxiv.org/abs/2201.11903)  |   |  
|   | Thu (3/19)  | [LLMs: Post-training](https://princeton-nlp.github.io/cos484/lectures/L12_sp26.pdf)  |  1. [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155) (InstructGPT)  
 |   |  
|   | Fri (3/20)  | [Precept 6](https://princeton-nlp.github.io/cos484/precepts/precept6_s26.pdf)  |   |   |  
| 9  | Tue (3/24)  | [LLMs: advanced techniques](https://princeton-nlp.github.io/cos484/lectures/L13_sp26.pdf)  |  1. [Direct Preference Optimization: Your Language Model is Secretly a Reward Model](https://arxiv.org/abs/2305.18290)  
2. [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://arxiv.org/abs/2405.14734)  
3. [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685)  | Project proposals due  |  
|   | Thu (3/26)  | [Language Agents](https://princeton-nlp.github.io/cos484/lectures/L14_sp26.pdf)  |  1. [ReAct: Synergizing Reasoning and Acting in Language Models](https://arxiv.org/abs/2210.03629) (ReAct)  
2. [Tree of Thoughts: Deliberate Problem Solving with Large Language Models](https://arxiv.org/abs/2305.10601) (Tree of Thoughts)  
3. [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770)  
4. [SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering](https://arxiv.org/abs/2405.15793)  
5. [Cognitive Architecture for Language Agents (CoALA) ](https://arxiv.org/abs/2309.02427)  |   |  
|   | Fri (3/27)  | Precept 7  |   |   |  
| 10  | Tue (3/31)  | [Systems for LLM training](https://princeton-nlp.github.io/cos484/lectures/L15_sp26.pdf)  |  1. [All the Transformer Math You Need to Know](https://jax-ml.github.io/scaling-book/transformers/)  
2. [The Ultra-Scale Playbook: Training LLMs on GPU Clusters](https://huggingface.co/spaces/nanotron/ultrascale-playbook?section=high_level_overview)  | A3 due, [A4 out](https://princeton-nlp.github.io/cos484/assignments/a4_sp26.pdf)  |  
|   | Thu (4/2)  | [Systems for LLM inference](https://princeton-nlp.github.io/cos484/lectures/L16_sp26.pdf)  |  1. [All About Transformer Inference](https://jax-ml.github.io/scaling-book/inference/)  
 |   |  
|   | Fri (4/3)  | Precept 8  |   |   |  
| 11  | Tue (4/7)  | Guest Lecture: [Lei Li](https://lileicc.github.io/) (CMU) - Toward Safer LLMs and AI Agents. Zoom only ([recording](https://princeton.zoom.us/rec/share/tceMjdRgAX2BeuBA1KlQHOgjq9fnvnE5n9gKGMt8VGvis-cl7-_sZE23tw8CjHh7.QogwsTOkEFp6yRTJ))  |   |   |  
|   | Thu (4/9)  | [Reasoning with LLMs](https://princeton-nlp.github.io/cos484/lectures/L18_sp26.pdf)  |   |   |  
|   | Fri (4/10)  | Precept 9  |   |   |  
| 12  | Tue (4/14)  | Guest lecture: [Xingyu Fu](https://zeyofu.github.io/) (Princeton). In class.  |   | A4 due  |  
|   | Thu (4/16)  | Guest lecture: [Ofir Press](https://ofir.io/) (Meta). Zoom only.  |   |   |  
|   | Fri (4/17)  | Precept 10  |   |   |  
| 13  | Tue (4/21)  | Project meetings (no class)  |   |   |  
|   | Thu (4/23)  | Project meetings (no class)  |   |   |  
| 14  | Tue (4/28)  | Reading period  |   |   |  
|   | Tue (5/5)  | Project poster presentations (2-4pm, Friend Center atrium)  |   |   |  
|   | Mon (5/12)  | Dean's date  |   | Final project report due  |  
Coursework
Assignments
All assignments are due at **11:59pm EST** on Tuesdays. We allow each student up to 4 free late days that you can use any time during the semester, with at most 3 late days per assignment. No assignment will be accepted more than 3 days after the deadline. We trust you to use these late days responsibly to help mitigate special circumstances. Late hours will round up to a full day (i.e. if you submitted 4 hours late, it will cost you 1 late day). Once you run out of late days, each additional day late will incur a fixed 10% deduction for the given assignment (e.g. 10 points off a 100-point assignment). For extenuating circumstances (e.g, those that involve a Dean), please make sure to email the Dean and the course instructors. For students with a dean’s note, the weight of their missed/penalized assignment will be added to the midterm and your midterm score will be scaled accordingly (for homeworks 0, 1 and 2) (e.g. if you are penalized 2 points overall, your midterm will be worth 27 and your score will be multiplied by 27/25). Missing homework 3 and 4 after the midterm can only be compensated by arranging an oral exam on the pertinent material. 
**Writeups:** Homeworks should be written up clearly and succinctly; you may lose points if your answers are unclear or unnecessarily complicated. Using LaTeX is recommended (here's a [template](https://princeton-nlp.github.io/cos484/assignments/template.tex)), but not a requirement. If you've never used LaTeX before, refer to this introductory guide on [Working with LaTeX](http://bit.ly/WorkingWithLaTeX) to get started. Hand-written assignments must be scanned and uploaded as a pdf. 
**Collaboration policy and honor code:** You are free to form study groups and discuss homeworks and projects. However, you must write up homeworks and code from scratch independently, and you must acknowledge in your submission all the students you discussed with. The following are considered to be honor code violations (in addition to the Princeton honor code): 
  * Looking at the writeup or code of another student.
  * Showing your writeup or code to another student.
  * Discussing homework problems in such detail that your solution (writeup or code) is almost identical to another student's answer.
  * Uploading your writeup or code (or released solutions) to a public repository (e.g. GitHub, Bitbucket, Pastebin) so that it can be accessed by other students. 

When debugging code together, you are only allowed to look at the input-output behavior of each other's programs (so you should write good test cases!). It is important to remember that even if you didn't copy but just gave another student your solution, you are still violating the honor code, so please be careful. If you feel like you made a mistake (it can happen, especially under time pressure!), please reach out to the instructors; the consequences will be much less severe than if we approach you. 
**Large language model policy:** You **may not** consult a Large Language Model (LLM) when working on the Theory portions of assignments. You may use coding assistants like GitHub Copilot/Cursor for completing the programming parts of the homework but we require you to document your use of these assistants in your submission (details will be provided in the assignment). You may also use LLMs to help understand concepts and to study for exams (e.g. like looking up YouTube videos explaining CNNs). However, note that LLMs can generate incorrect text and that they often are equally confident when they are correct vs. incorrect (this is analogous to the fact that not everything presented on YouTube is factually correct). Our course materials are the only standard with which exams and assessments will be evaluated. 
Final Project
[[Project guidelines are available here]](https://docs.google.com/document/d/e/2PACX-1vRhaTW6y1hhiPpbwrBF2lB4qv_LuavSuZfmhMOTqdgtZvJJwGbKcePwVri4XKMma2cJg1MSpWpfp-d9/pub)   
  
The final project offers you the chance to apply your newly acquired skills towards an in-depth NLP application. Students are required to complete the final project in teams of **3** students.   
  
There are **two** options for the final project: (a) reproducing an ACL/NAACL/EMNLP/COLM, or NLP papers from ICLR/ICML/NeurIPS from past 5 years (encouraged); (b) complete a research project (for this option, you need to discuss your proposal and get prior approval from an instructor). All the final projects will be completed in teams of 3 students (Find your teammates early!). 
**Deliverables** : The final project is worth 35% of your course grade. The deliverables include: 
  * **Proposal** (0%): You need to turn in a one-page proposal on **March 20th**. The proposal should outline what you propose to do and a rough plan for how you will pursue the project. We will then provide feedback and guidance on the direction to maximize the project’s chance of succeeding. _This proposal is not graded._
  * **Project presentation** (10%): At the end of the semester, we will schedule project presentations for all the projects in the class.
  * **Final paper** (25%): You need to complete a final report in the style of a conference submission (we recommend you to use the [ACL 2023 template](https://princeton-nlp.github.io/cos484/assignments/acl2023.zip)). It should begin with an abstract and introduction, clearly describe the proposed idea or exploration, present technical details, give results, compare to baselines, provide analysis and discussion of the results, and cite any sources you used. 


**Policy and honor code:**
  * The final projects are required to be implemented in Python. You can use any deep learning framework such as PyTorch and Tensorflow. 
  * You are free to discuss ideas and implementation details with other teams. However, under no circumstances may you look at another team's code, or incorporate their code into your project. 
  * Do not share your code publicly (e.g. in a public GitHub repo) until after the class has finished.


Submission
**Electronic Submission:** Assignments and project proposal/paper are to be submitted as pdf files through Gradescope. If you need to sign up for a Gradescope account, please use your @princeton.edu email address. You can submit as many times as you'd like until the deadline: we will only grade the last submission. **Submit early to make sure your submission uploads/runs properly on the Gradescope servers.** If anything goes wrong, please ask a question on Ed or contact a TA. Do not email us your submission. **Partial work is better than not submitting any work.** For more detailed information on submitting your assignment solutions, see this guide on [assignment submission logistics](http://bit.ly/COS_NLP_Submission). 
For assignments with a programming component, we may automatically sanity check your code with some basic test cases, but we will grade your code on additional test cases. **Important** : just because you pass the basic test cases, you are by no means guaranteed to get full credit on the other, hidden test cases, so you should test the program more thoroughly yourself! 
**Regrades** : If you believe that the course staff made an objective error in grading, then you may submit a regrade request. Remember that even if the grading seems harsh to you, the same rubric was used for everyone for fairness, so this is not sufficient justification for a regrade. It is also helpful to cross-check your answer against the released solutions. If you still choose to submit a regrade request, click the corresponding question on Gradescope, then click the "Request Regrade" button at the bottom. Any requests submitted over email or in person will be ignored. Regrade requests for a particular assignment are due one week after the grades are returned. Note that we may regrade your entire submission, so depending on your submission you may actually lose more points than you gain. 
FAQ
  * **Q: Can we do the project in a team of 2?**   
A: We strongly encourage you to have 3 members on your team. If you really want work in a team of 2, please write to the course staff on Ed with a justification and a rough plan of the project. We want to make sure the scope and workload of the project is reasonable and note that we will grade projects regardless of team sizes. 
  


  

