# Info 159/259. Natural Language Processing

https://people.ischool.berkeley.edu/~dbamman/nlp25.html

# [Natural Language Processing](https://people.ischool.berkeley.edu/~dbamman/index.html)
Info 159/259. Spring 2025  
T/Th 3:30-5pm, Lewis 100  
[David Bamman](http://people.ischool.berkeley.edu/~dbamman/) (dbamman@berkeley.edu)   

## Info
This course introduces students to natural language processing and exposes them to the variety of methods available for reasoning about text in computational systems. NLP is deeply interdisciplinary, drawing on both linguistics and computer science, and helps drive much contemporary work in text analysis (as used in computational social science, the digital humanities, and computational journalism). We will focus on major algorithms used in NLP for various applications (part-of-speech tagging, parsing, coreference resolution, machine translation) and on the linguistic phenomena those algorithms attempt to model. Students will implement algorithms and create linguistically annotated data on which those algorithms depend. 
## Staff
[David Bamman](http://people.ischool.berkeley.edu/~dbamman/) (dbamman@berkeley.edu), OH: Wed 10am-noon (312 South Hall). **TAs (info159259-instructors@lists.berkeley.edu)** : 
* Hellina Hailu Nigatu (hellina_nigatu@berkeley.edu)
* Lucy Li (lucy3_li@berkeley.edu)
* Aaron Rodden (aaron_rodden@berkeley.edu)
* Mackenzie Cramer (mackenzie.hanh@berkeley.edu)
* Zikai Liu (zikailiu@berkeley.edu)
## TA Office Hours
* Monday: 11-12pm, South Hall 210
* Wednesday: 1-2pm, South Hall 202
* Friday: 12-1pm, South Hall 210
## Texts
  * [SLP3] Dan Jurafsky and James Martin, Speech and Language Processing (3nd ed. draft) [Available [here](https://web.stanford.edu/~jurafsky/slp3/)]
  * [PS] James Pustejovsky and Amber Stubbs, Natural Language Annotation for Machine Learning (2012) [Online access available for free through the UC library. [[Library link](https://search.library.berkeley.edu/discovery/fulldisplay?docid=alma9914828954906531&context=L&vid=01UCS_BER:UCB&lang=en&search_scope=DN_and_CI&adaptor=Local%20Search%20Engine&tab=Default_UCLibrarySearch&query=any,contains,Natural%20Language%20Annotation%20for%20Machine%20Learning%20by%20James%20Pustejovsky,%20Amber%20Stubbs&offset=0)] [[O'Reilly link](https://learning.oreilly.com/library/view/natural-language-annotation/9781449332693/ch06.html)]


## Syllabus
(Subject to change.)  
| Week  | Date  | Topic  | Required  | Optional  |  
| --- | --- | --- | --- | --- |  
| 1  | 1/21  | Intro  |   |   |  
| 1/23  | Words  | [SLP2](https://web.stanford.edu/~jurafsky/slp3/2.pdf)  | [SLP22](https://web.stanford.edu/~jurafsky/slp3/22.pdf)  |  
| 2  | 1/28  | Lexical semantics/word embeddings  | [SLP6](https://web.stanford.edu/~jurafsky/slp3/6.pdf)  |   |  
| 1/30  | Text classification  | [SLP5](https://web.stanford.edu/~jurafsky/slp3/5.pdf)  | [SLP4](https://web.stanford.edu/~jurafsky/slp3/4.pdf)  |  
| 3  | 2/4  | Annotation  | [PS ch. 6](https://learning.oreilly.com/library/view/natural-language-annotation/9781449332693/ch06.html)  |   |  
| 2/6  | Neural models for classification  | [SLP7](https://web.stanford.edu/~jurafsky/slp3/7.pdf)  |   |  
| 4  | 2/11  | Attention/Transformers  | [SLP9](https://web.stanford.edu/~jurafsky/slp3/9.pdf)  |   |  
| 2/13  | Language models 1  | [SLP3](https://web.stanford.edu/~jurafsky/slp3/3.pdf)  |   |  
| 5  | 2/18  | Masked language models  | [SLP11](https://web.stanford.edu/~jurafsky/slp3/11.pdf)  | [SLP8](https://web.stanford.edu/~jurafsky/slp3/8.pdf)  |  
| 2/20  | In-class Midterm 1  |   |   |  
| 6  | 2/25  | LLMs 1  | [SLP10](https://web.stanford.edu/~jurafsky/slp3/10.pdf)  |   |  
| 2/27  | LLMs 2  | [SLP10](https://web.stanford.edu/~jurafsky/slp3/10.pdf)  |   |  
| 7  | 3/4  | LLMs 3  | [SLP12](https://web.stanford.edu/~jurafsky/slp3/12.pdf)  |   |  
| 3/6  | Question answering and RAG  | [SLP14](https://web.stanford.edu/~jurafsky/slp3/14.pdf)  |   |  
| 8  | 3/11  | LLM agents  |   |  [Wang et al. 2024](https://arxiv.org/abs/2308.11432); [Park et al. 2023](https://arxiv.org/abs/2304.03442)  |  
| 3/13  | NER + sequence labeling  | [SLP17](https://web.stanford.edu/~jurafsky/slp3/17.pdf)  |   |  
| 9  | 3/18  | Dependency parsing  | [SLP19](https://web.stanford.edu/~jurafsky/slp3/19.pdf)  |   |  
| 3/20  | Word sense disambiguation  | [SLPG](https://web.stanford.edu/~jurafsky/slp3/G.pdf)  |   |  
| 10  | 3/25  | Spring break (no class)  |   |   |  
| 3/27  | Spring break (no class)  |   |   |  
| 11  | 4/1  | Information extraction  | [SLP20](https://web.stanford.edu/~jurafsky/slp3/20.pdf)  |   |  
| 4/3  | In-class Midterm 2  |   |   |  
| 12  | 4/8  | Machine translation and multilingual NLP  | [SLP13](https://web.stanford.edu/~jurafsky/slp3/13.pdf)  |   |  
| 4/10  | Dialogue  | [SLP15](https://web.stanford.edu/~jurafsky/slp3/15.pdf)  |   |  
| 13  | 4/15  | Coreference resolution  | [SLP23](https://web.stanford.edu/~jurafsky/slp3/23.pdf)  |   |  
| 4/17  | Unsupervised models  | [Blei et al. 2012](https://www.cs.columbia.edu/~blei/papers/Blei2012.pdf)  |   |  
| 14  | 4/22  | Social NLP  |  [Nguyen et al. 2020](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2020.00062/full); [Ziems et al. 2024](https://aclanthology.org/2024.cl-1.8/)  |  [Olteanu et al. 2019](https://www.frontiersin.org/journals/big-data/articles/10.3389/fdata.2019.00013/full); [Wallach et al. 2018](https://cacm.acm.org/opinion/computational-social-science-computer-science-social-data/)  |  
| 4/24  | Vision/language models  |   |   |  
| 15  | 4/29  | Project presentations  |   |   |  
| 5/1  | Project presentations  |   |   |  
## Prerequisites
  * — Algorithms: Computer Science 61B
  * — Probability/Statistics: Computer Science 70, Math 55, Statistics 134, Statistics 140 or Data 100
  * — Strong programming skills


## Colab
Most assignments will be in the form of Jupyter notebooks that can be run locally on your computer or on Google Colab. Some assignments (especially those involving LLMs) will require access to a GPU; you can expect to require Colab Pro-level ($10) access for about two months. 
## Grading
#### Info 159  
| 25%  | Homeworks (3 late days)  |  
| --- | --- |  
| 10%  | Quizzes  |  
| 20%  | Midterm exams  |  
| 25%  | Annotation project (AP)  |  
| 20%  | NLP subfield survey  |  
#### Info 259  
|  25%  | Homeworks (3 late days)  |  
| --- | --- |  
| 10%  | Quizzes  |  
| 20%  | Midterm exams  |  
| 45%  | Project:  |  
|   |  5% Proposal/literature review  |  
|   |  10% Midterm report  |  
|   |  25% Final report  |  
|   |  5% Presentation  |  
Lectures will be recorded through course capture and made available through bCourses. 
## Quizzes
We will have frequent in-class pop quizzes (12-15 over the entire semester). The topics of the quizzes will be drawn from the required readings assigned for that day, along with any lecture content from that class period. Be sure to do the required readings before class! If you miss class for any reason, you will not be able to make up a quiz. In order to give you some flexibility for necessary absences, we will drop your three lowest-scoring quizzes over the semester when calculating your final quiz grade. 
## Midterms
This course has two midterm exams scheduled for 2/20 and 4/3 (completed in-person during class time) and no final exam. If you cannot take a midterm due to an illness or other medical reason (not travel or other conflicts), email us (info159259-instructors@lists.berkeley.edu) before the exam. Your midterm exam grade for the course will be the average of the two midterms.
## Annotation project
The most exciting applications of NLP haven't been invented yet. While much of this course will give you exposure to the common _methods_ in NLP, you will also carry out an annotation project where you will get exposure to the entire NLP design process for building a classifer for a brand new task. You will decide on a new document classification NLP task, annotate data to support it (including creating annotation guidelines), measure your inter-annotator agreement rate, and build a classifier to predict those labels using the methods we discuss in class.
You may use existing NLP tasks (except **sentiment analysis**), but try to think outside of the box: projects will be rewarded for their **creativity and originality** in coming up with a task that few people have considered before, and for being able to create a comprehensive set of guidelines that lead to consistent third-party annotations (i.e., not by your team). To give you a sample of similarly creative new NLP tasks, consider the following work: [how dogmatic is a forum post?](https://arxiv.org/pdf/1609.00425.pdf); [how respectful are police officers in their interactions at traffic stops?](http://www.pnas.org/content/114/25/6521.full.pdf); [how suspenseful is a passage from a story?](http://markalgeehewitt.org/index.php/main-page/projects/the-machinery-of-suspense/); [how much time is passing in it?](https://www.ideals.illinois.edu/bitstream/handle/2142/91604/WhyLiteraryTimeIsMeasuredInMinutes.pdf?sequence=2)
#### Deliverables
(These are a summary of the deliverables; see bCourses for the official description of them.)
* AP1. Form a project group of exactly 4 people and let us know who's in the group. Either select your group yourself or let us pair you randomly with other teammates. Tell us in 100 words the basic idea, so we can give you feedback early.
* AP2. Decide on a document classification annotation task. No sentiment analysis! Collect data and tokenize it. All data must be shareable with the public, so no private information, nothing within copyright (e.g., no song lyrics), and nothing from social media (including Reddit). Keep privacy and ethics in mind as you are considering potential sources of data.
* AP3. Create a robust set of annotation guidelines that describe the boundaries of your categories and how you should annotate your data. Use those guidelines to annotate a small sample of 50 documents. 
* AP4. Annotate your data independently using those guidelines. (This is an individual assignment.)
* AP5. Build a classifier to automatically predict the labels using the data you've annotated. 
* AP6. Present your work to the class! 
## NLP subfield survey (Info 159)
Understanding how to read and synthesize articles in NLP is an important part of carrying out research in this space. To cultivate this skill, your final report will be a 2000-word survey for a specific NLP subfield of your choice (e.g., coreference resolution, question answering, interpretability, narrative generation, etc.), synthesizing at least 25 papers published at ACL, EMNLP, NAACL, EACL, AACL, _Transactions of the ACL_ or _Computational Linguistics_. This survey should be able to provide a newcomer (such as yourself at the start of the semester) a sense of the current state of the art in that subfield in 2025, the major historical papers that have defined that area, and the different schools of thought within it. The survey should use the ACL style files for formatting, which are available as an [Overleaf template](https://www.overleaf.com/latex/templates/association-for-computational-linguistics-acl-conference/jvxskxpnznfj). 
## Grad Project (Info 259)
Info 259 will be capped by a semester-long project (involving one to three students), involving natural language processing -- either focusing on core NLP methods or using NLP in support of an empirical research question. For examples of the former, see papers published at [ACL](https://www.aclweb.org/anthology/venues/acl/), [NAACL](https://www.aclweb.org/anthology/venues/naacl/) and [EMNLP](https://www.aclweb.org/anthology/venues/emnlp/); for examples of the latter, see workshops for [NLP and Computational Social Science](https://www.aclweb.org/anthology/venues/nlpcss/), [Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature](https://www.aclweb.org/anthology/venues/latech/), [Natural Language Processing Techniques for Educational Applications](https://www.aclweb.org/anthology/venues/nlptea/), [Noisy User-Generated Text](https://www.aclweb.org/anthology/venues/wnut/), and [many more](https://www.aclweb.org/anthology/venues/). 
The project will be comprised of four components:
  * — Project proposal and literature review. Students will propose the research question to be examined, motivate its rationale as an interesting question worth asking, and assess its potential to contribute new knowledge by situating it within related literature in the scientific community. (1000 words; 5 sources)
  * — Midterm report. By the middle of the course, students should present initial experimental results and establish a validation strategy to be performed at the end of experimentation. (2000 words; 10 sources)
  * — Final report. The final report will include a complete description of work undertaken for the project, including data collection, development of methods, experimental details (complete enough for replication), comparison with past work, and a thorough analysis. Projects will be evaluated according to standards for conference publication—including clarity, originality, soundness, substance, evaluation, meaningful comparison, and impact (of ideas, software, and/or datasets). (3000 words, not including references)
  * — Presentation. At the end of the semester, teams will present their work to the class.

All reports should use the ACL style files for formatting, which are available as an [Overleaf template](https://www.overleaf.com/latex/templates/association-for-computational-linguistics-acl-conference/jvxskxpnznfj). 
## Policies
### Academic Integrity
All students will follow the UC Berkeley [code of conduct](http://sa.berkeley.edu/code-of-conduct). You may discuss homeworks at a high level with your classmates (if you do, include their names on the submission), but each homework deliverable must be completed independently -- all writing and code must be your own. All quizzes and exams must be completed on your own. If you mention the work of others, you must be clear in citing the appropriate source (For additional information on plagiarism, see [here](http://gsi.berkeley.edu/gsi-guide-contents/academic-misconduct-intro/plagiarism/) and [this great infographic by Emily Myers](https://people.ischool.berkeley.edu/~dbamman/Myers-Plagiarism-Infographic.pdf).) This holds for source code as well: if you use others' code (e.g., from StackOverflow), you must cite its source. All homeworks and project deliverables are due at the time and date of the deadline. We have zero tolerance policy for cheating and plagiarism; violations will be referred to the Center for Student Conduct and will likely result in failing the class.
### AI Assistants
This is a class on NLP, and LLMs are NLP technologies; one goal of this course is to understand how to use LLMs sensibly while still prioritizing learning outcomes. So what are the learning outcomes of this class?
  * — Above all, you are cultivating broad knowledge of the landscape of methods available to use in NLP for answering questions involving text as data, and developing an understanding of what methods are appropriate to use for a given task. When given a new task that no one has seen before, you should be able to know how to think about it and decide what methods are good and which ones are bad. If you ask an LLM "what should I do?", you need to be able to assess the quality of the response. You are cultivating a sense of taste that comes from experience, so you need to be sure to give yourself that experience thinking through problems.
  * — You are getting experience implementing those ideas in code. AI Assistants (e.g., Copilot, etc.) can dramatically help with this and provide a natural sandbox to implement ideas and see if they work. A learning outcome for this class is not to teach you Python, so I don't care if you're using Copilot to tell you how to read in a CSV file or sort a Python dict by value. But you should understand what it's doing, be able to answer questions about it and be empowered to edit it to do something different. You are always ultimately responsible for the code you write.
  * — You will read a range of papers in NLP to see how creative people have been in applying it to measurement problems. The goal here (as noted above) is to both give you a sense of the landscape of research that you will be working within (many people have thought about your ideas before, so you should see what they have found), but also how research is structured -- e.g., how research questions, models, and validation all fit together. You are again cultivating your sense of research taste as you read these papers, and this is only developed through lots of experience and reading. 


I _highly_ recommend that you don't try to game the homeworks and subfield survey with LLMs. They're not there as busywork, but to give you repeated experience cultivating the tastes mentioned above. Work through them. Feel free to use automatic coding assistants for some low-level aspects of this work (as you'd certainly do at your job/research in the future), but be sure to work through the core of the challenge yourself in order to get that experience solving these kinds of problems. We'll see how you've been able to cultivate that taste over the course of the semester; but consider beyond this class in the future -- as you talk with people who work in NLP (at conferences or job interviews), will your experience come through in those moments without ChatGPT to help you? 
To make sure we have a clean dividing line between fair and unfair use of automatic writing/coding assistants, you retain ultimate responsibility the behavior of any automatic assistant you use. If an LLM plagiarizes the text that someone else wrote (and you then submit that text as your own), you're the one who's ultimately accountable for plagiarizing. Even if assisted by code suggestions, any work you submit should still be the product of your own creation, representing your own ideas, code and words, and submissions that rely too heavily on AI tools will not be graded very favorably. You are free to use LLMs to rephrase text that you are the author of, or to help brainstorm ideas, but you must be the author of any text you submit (e.g., homeworks, subfield survey or project reports), and you cannot simply be editing the text that originated in an LLM. You must know what all code you submit does; if we ask and you don't know, your grade will be influenced by it.
I hope this goes without saying, but for the Annotation Project, you absolutely cannot use any automatic method (including LLMs) to label your data for you. The goal of that project is for you to come up with an original task and create the labeled data to support it using your own expert judgment. You are creating a benchmark that others can trust to use when evaluating automatic methods, so it cannot be created by those methods themselves.
Likewise, you should not use LLMs to create any part of your subfield survey for you. You are free to use LLMs as you would a search engine (e.g., to get pointers to some relevant papers to get you on an information seeking trail), but do not use them to substitute for your own careful reading, literature review, and synthesis. As we will discuss in class, LLMs are prone to hallucination, factual errors, and other forms of bias. Fabricated citations will likely lead to a failing grade. As noted above, you must be the originating author of any text you submit. 
For the use of any AI assistance in your final project deliverables, you should follow the [ACL 2023 Policy on AI Writing Assistance](https://2023.aclweb.org/blog/ACL-2023-policy/). 
### Ed Discussion
We'll use Ed Discussion as a platform for asking and answering questions about the course material, including homeworks. Students are encouraged to actively participate on this forum and help others by answering questions that arise (helpful students can see a grade bump across a threshold (e.g., B+ to A-) for this participation. When helping with homework questions, keep the discussion to the high-level concepts; don't post answers to homeworks or quiz/exam questions.
### TA office hours
While at TA office hours, keep academic integrity in mind: you may discuss homework questions at a high level with others present, but don't discuss specific answers or share screens with code solutions. Neither the TA office hours nor Ed Discussion should be used for pre-grading (asking if a specific answer to a homework or quiz question is correct before the assignment is due).
### Students with Disabilities
Our goal is to make class a learning environment accessible to all students. If you need disability-related accommodations and have a Letter of Accommodation from the DSP, have emergency medical information you wish to share with me, or need special arrangements in case the building must be evacuated, please inform me immediately. I'm happy to discuss privately after class or at my office.
### Late assignments
Student have will have a total of **three** late days to use when turning in homework assignments (not group annotation project deliverables, 259 project deliverables, or the subfield survey); each late day extends the deadline by 24 hours. If all late days have been used up, homeworks can be turned in up to 48 hours late for 50% credit; anything submitted after 48 hours late = 0 credit. Each homework will be due at 11:59pm, and will have a 2-hour grace period for any last-minute submission issues. Late days and incompletes will be assessed immediately following the grace period (at 2:00am sharp). The grace period applies to late days as well (if a homework is due at 11:59pm 1/21, and you use a late day to extend it to 11:59pm 1/22, you may turn it in up to 2:00am 1/23 and still be assessed 1 late day.) Late days are assessed immediately once homeworks are submitted late and can't be retroactively changed (if you submit 4 homeworks late, for example, you can't decide after the fact which ones to apply your 3 slip days to -- they apply to whichever homeworks use them up first).
### Curving
Grades for this course will **not** be curved. Minimum thresholds for letter grades are the following: 93 A, 90 A-, 87 B+, 83 B, 80 B-, 77 C+ 73 C, 70 C-, 67 D+, 63 D, 60 D-, 0 F. Students taking the course P/NP must complete all deliverables and will receive a P if their grade is greater or equal to 70 (C-); Students taking S/U must complete all deliverables and will receive an S if their grade is greater or equal to 80 (B-).
