# Advances in Computer Vision – Scene Representation Group

https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#syllabus

[Skip to content](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#main-content) [ Scene Representation Group ](https://scenerepresentations.org/) [Publications](https://scenerepresentations.org/publications) [Talks](https://scenerepresentations.org/talks) [Teaching](https://scenerepresentations.org/courses) [People](https://scenerepresentations.org/people)
[ Massachusetts Institute of Technology ](https://web.mit.edu "The Massachusetts Institute of Technology") [ ](https://www.csail.mit.edu/ "MIT Computer Science & Artificial Intelligence Lab")
  1. [Teaching](https://scenerepresentations.org/courses)
  2. Advances in Computer Vision 6.8300 spring 2026

Advances in Computer Vision - MIT 6.8300
  * [Details](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#details)
  * [Syllabus](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#syllabus)
  * [Related Courses](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#related-courses)
  * [Gradescope](https://www.gradescope.com/courses/1241394)
  * [Slides](https://drive.google.com/drive/folders/1glgfywpMzX4i9YM6Ko60GFz4Li5fmmL7?usp=sharing)
  * [Piazza](https://piazza.com/class/mkvf5osqna46qo/)
  * [Recordings](https://mit.hosted.panopto.com/Panopto/Pages/Sessions/List.aspx?folderID=4b7e677f-0750-4ac6-9ae1-b3d0011f240b)


### Course Contents
This course dives into advanced concepts in computer vision. A first focus is geometry in computer vision, including image formation, a closer look at the fourier transform and its relationship to geometric deep learning, classic multi-view geometry, multi-view geometry in the age of deep learning, differentiable rendering, neural scene representations, correspondence estimation, optical flow computation, and point tracking.
Next, we explore generative modeling and representation learning including image and video generation, guidance in diffusion models, conditional probabilistic models, as well as representation learning in the form of contrastive and masking-based methods. 
Finally, we will explore the intersection of robotics and computer vision with imitation learning and world models.
### Prerequisites
The formal prereqs of this course are: 6.7960 Deep Learning, (6.1200 or 6.3700), (18.06 or 18.C06).
This class is an advanced graduate-level class. You have to have working knowledge of the following topics, i.e., be able to work with them in numpy / scipy / pytorch. There will be no explainer on this and TAs will not be able to help you with these basics.
Deep Learning: Proficiency in Python, Numpy, and PyTorch, vectorized programming, and training deep neural networks. Convolutional neural networks, transformers, MLPs, backpropagation.
Linear Algebra: Vector spaces, matrix-matrix products, matrix-vector products, change-of-basis, inner products and norms, Eigenvalues, Eigenvectors, Singular Value Decomposition, Fourier Transform, Convolution.
### Schedule
6.8300 will be held as 1.5 hour long lectures in room **26-100** :   
|  **Tuesday**  |  1:00 – 2:30 pm  |  
| --- | --- |  
|  **Thursday**  |  1:00 – 2:30 pm  |  
### Collaboration Policy
Problem sets should be written up individually and should reflect your own individual work. You cannot copy code from another student. However, you may discuss with your peers, TAs, and instructors.
You should not copy or share complete solutions or ask others if your answer is correct, whether in person or via Piazza or Canvas.
If you work on the problem set with anyone other than TAs and instructors, list their names at the top of the problem set.
### Office Hours  
|  **Monday**  |  2:00 – 3:00 pm  |  36-156   |  
| --- | --- | --- |  
|  **Tuesday**  |  5:00 – 6:00 pm  |  [Zoom](https://mit.zoom.us/j/91698689393)  |  
|  **Wednesday**  |  11:00 – 12:00 pm  |  36-112   |  
|  **Thursday**  |  5:00 – 6:00 pm  |  [Zoom](https://mit.zoom.us/j/91698689393)  |  
|  **Friday**  |  3:00 – 4:00 pm  |  36-153   |  
Office Hours are subject to change based on staff availability. Please check Piazza for Office Hour schedule updates.
### Late Submissions Policy
There are no late days for problem sets, the submissions close with the deadline and late submissions will not be graded.
### Grading Policy
Grading will be split between five module-specific problem sets and a final project:   
|  10%   |  **Problem Sets**  
5 problem sets note our separate policies on [Collaboration](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#collaboration-policy), [AI Assistants](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#ai-policy), and [Late Submissions](https://scenerepresentations.org/courses/2026/spring/advances-in-cv/#late-policy).  |  
| --- | --- |  
|  45%   |  **In-class Midterm**  
Pen-and-paper, closed-book in-class quiz. Will deal with content of homework assignments and lectures.  |  
|  45%   |  **Final Project**  
Blog Post (80%) + Recorded two-minute Talk(20%)  |  
The final project will be a **research project on perception** of your choice: 
  * You will run experiments and do analysis to explore your research question. 
  * You will write up your research in the format of a blog post. Your post will include an explanation of background material, new investigations, and results you found. 
  * You are encouraged to include plots, animations, and interactive graphics to make your findings clear. [Here](https://distill.pub/) [are](https://www.engraved.blog/why-machine-learning-algorithms-are-hard-to-tune/) [some](http://karpathy.github.io/2015/05/21/rnn-effectiveness/) [examples](https://ai.facebook.com/blog/dino-paws-computer-vision-with-self-supervised-transformers-and-10x-more-efficient-training) of well-presented research. 

The final project will be graded for clarity and insight as well as novelty and depth of the experiments and analysis. Detailed guidance will be given later in the semester. 
### FAQ  
| Q  | Can I take this course if I have _not_ taken 6.7960 Deep Learning or a comparable class?  |  
| --- | --- |  
| A  | We advise against it. There will be homework assignments where you will be asked to re-implement deep learning papers by yourself. If you don't have working knowledge in Deep Learning using Pytorch, you are unlikely to perform well on these assignments. We will generally not discuss topics that were discussed in the Deep Learning class, i.e., we will not be reiterating Transformers, CNNs, how to train these models, etc, but will assume that you are already familiar with them.  |  
| Q  | Is this class a CI-M class?  |  
| A  | No, this is a graduate class.  |  
| Q  | Is 6.8301 (the undergraduate version) taught this semester?  |  
| A  | The undergraduate version is taught this semester as well. For logistical reasons, it had to be renamed to [6.S058 / 6.4300, and is taught by Profs. Bill Freeman and Phillip Isola](https://introtocv.github.io/). 6.S058 is a CI-M class, and does _not_ have a prerequisite on 6.7960 Deep Learning.  |  
| Q  | Is attendance required? Will lectures be recorded?  |  
| A  | Attendance is at your discretion. Yes, lectures will be recorded and uploaded.  |  
| Q  | Is this a TQE course?  |  
| A  | Yes.  |  
### Contact Instructors
If you have a question regarding extensions, accommodations, the midterm quiz, or other logistics, please email 6.830-instructors-sp26@mit.edu. Please note that extensions will ONLY be offered for extenuating circumstances with S³ support, so please reach out to S³ first and cc your S³ dean when you reach out to us.
### Instructors
Frédo Durand, Vincent Sitzmann, Peter E. Holderrieth
### AI Assistants Policy
Different from last year, this year, we welcome you to finish the problem set with the help of AI assistants. The homeworks are designed to give you a deeper, practical understanding of the course material, but are not the primary means of assessment any more - the midterm quiz, which will deal with content of homework assignments and lectures, will be the primary means of assessment. We used this additional degree of freedom to make the homework assignments more educational and interactive, and it will be easier to judge at the time of submission if you got everything right.
## Syllabus  
| 
###  Module 0: Introduction to Computer Vision
 |  
| --- |  
| 
#### Introduction to Vision
Tue, Feb. 3rd  | 
  * Administrativa & Logistics 
  * Historical perspective on vision: problems identified so far 
  * What is vision? 
  * Outlook 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=76432f2a-ed9f-4b03-9223-b3df01002fd3)
  * [Slides ](https://drive.google.com/file/d/1zkM0mBOT8gu7u5ngT-uTAUtVkY7215Pq/view?usp=drive_link)

 |  
| 
###  Module 1: Module 1: Geometry, 3D and 4D
 |  
| 
#### What is an Image: Pinhole Cameras & Projective Geometry
Thu, Feb. 5th  | 
  * Image as a 2D signal 
  * Image as measurements of a 3D light field 
  * Pinhole camera and perspective projection 
  * Camera motion and poses 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=d246c024-0687-48a9-af7d-b3df01003041)
  * [Slides ](https://drive.google.com/file/d/1T3dbXri-7IZppY44kLIwQnpyiS4p3iEx/view?usp=drive_link)


  * [ pset 1 out ](https://drive.google.com/file/d/17rmXoNZopCC24sCSibmJcupSNAkBRojw/view?usp=sharing)

 |  
| 
#### Linear Image Processing & Transformations
Tue, Feb. 10th  | 
  * Images as functions: continuous vs discrete 
  * Function spaces and Fourier transform overview 
  * Image filtering: gradients, Laplacians, convolutions 
  * Multi-scale processing: Laplacian and multi-scale pyramids 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=915ae91b-f038-4ae8-b4a7-b3df01003074)
  * [Slides ](https://drive.google.com/file/d/1s6K0n9YO2fM7su5KkMjWtzbswSSRcXJ2/view?usp=drive_link)

 |  
| 
#### Representation Theory in Vision
Thu, Feb. 12th  | 
  * Groups 
  * Group Representations 
  * Steerable Bases 
  * Invariant Operators 
  * Finding Steerable Bases via the Eigendecomposition of Invariant Operators 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=9f8edd66-ffad-4b43-b736-b3df01003090)
  * [Slides ](https://drive.google.com/file/d/1tHcOg9V9xAEZ_FcdE1KqAGjNAvdsX9sw/view?usp=drive_link)

 |  
| 
#### No Class (Monday Schedule)
Tue, Feb. 17th holiday  
 |   | 
  * pset 1 due

  

  * [ pset 2 out ](https://github.com/6-8300/pset-2-2026)

 |  
| 
#### Geometric Deep Learning and Vision
Thu, Feb. 19th  | 
  * Equivariance and invariance 
  * Regular Group Convolutions 
  * Steerable Group Convolutions 
  * Challenges of applying geometric techniques to vision tasks 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=647f2ae6-40cb-4037-9b13-b3df010030de)
  * [Slides ](https://drive.google.com/file/d/1yg7OBMV9FHor5U4e9Tu7p1_KCM4p54qM/view?usp=drive_link)

 |  
| 
#### Snow day
Tue, Feb. 24th holiday  
 |   |   |  
| 
#### Optical Flow
Thu, Feb. 26th  | 
  * What is optical flow? 
  * Color Constancy Assumption 
  * Infinitesimal Optical Flow 
  * Multi-Scale Cost and Correlation Volumes 
  * Learning-based optical flow 
  * RAFT 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=ae39a3c3-df98-4833-a54a-b3df01003123)
  * [Slides ](https://drive.google.com/file/d/1d9W7VuLRx1Bnfvs0NHpgAiorrUtBej5x/view?usp=drive_link)


  * pset 2 due

 |  
| 
#### Point Tracking, Scene Flow and Feature Matching
Tue, March 3rd  | 
  * Point Tracking 
  * Scene Flow 
  * Connection of Scene Flow and Pixel Motion, FlowMap 
  * Sparse Correspondence and Invariant Descriptors 
  * SIFT 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=d2be2add-0f20-463e-8a47-b3df0100313f)
  * [Slides ](https://drive.google.com/file/d/1HnG0lijTPp6pOO6h6fEoJy_fPqFD1QKU/view?usp=drive_link)


  * [ pset 3 out ](https://github.com/6-8300/pset3-2026)

 |  
| 
#### Multi-View Geometry
Thu, March 5th  | 
  * Triangulation in Light Fields: Infinetismal perspective 
  * Finite Triangulation 
  * Epipolar Geometry 
  * Eight-point algorithm and bundle adjustment 
  * Learning-Based Approaches: Dust3r & Mast3r 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=4c7465f1-db04-4658-8998-b3df0100315f)
  * [Slides ](https://drive.google.com/file/d/1X5j93g5UsK4KT_Zyld1UA96wDUvmKxLs/view?usp=drive_link)

 |  
| 
#### Differentiable Rendering: Data Structures and Signal Parameterizations
Tue, March 10th  | 
  * Surface-Based Representations 
  * (Volumetric) Field Representations 
  * Grid-based and adaptive data structures 
  * Neural Fields 
  * Hybrid Neural / discrete fields 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=f2ffd746-04e8-4e07-9528-b3df0100317d)
  * [Slides ](https://drive.google.com/file/d/1e3Fteq1lTm-JJj58MV12tLpUebTYdh4z/view?usp=drive_link)


  * pset 4 out

 |  
| 
#### Differentiable Rendering: Novel View Synthesis
Thu, March 12th  | 
  * Sphere tracing and volume rendering 
  * Differentiable rendering techniques 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=b7a6553e-e161-4d9a-9381-b3df01003199)
  * [Slides ](https://drive.google.com/file/d/1EFeC7R_j-wZtJVnlyQeHDzDGxux3m0JK/view?usp=drive_link)


  * pset 3 due

 |  
| 
#### Differentiable Rendering: Novel View Synthesis 2
Tue, March 17th  | 
  * Gaussian splatting 
  * Advanced differentiable rendering methods 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=ec9a43a3-29df-4807-a158-b3df010031bc)
  * [Slides ](https://drive.google.com/file/d/1VWHP82dI_mKHmNKg0HvjtVIViHSj3Fmp/view?usp=drive_link)


  * Final Project Guidelines Released

 |  
| 
#### Differentiable Rendering: Prior-Based 3D Reconstruction and Novel View Synthesis
Thu, March 19th  | 
  * Global inference techniques 
  * Light field inference and generative models 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=bdb6cb0b-76e8-4f41-ae59-b3df010031d7)
  * [Slides ](https://drive.google.com/file/d/17ePHnWwxEYescS38lajDcGpEV6ierN1j/view?usp=drive_link)


  * pset 4 due

 |  
| 
#### Student Holiday: Spring Break
Tue, March 24th holiday  
 |   |   |  
| 
#### Student Holiday: Spring Break
Thu, March 26th holiday  
 |   |   |  
| 
###  Module 2: Module 2: Unsupervised Representation Learning and Generative Modeling
 |  
| 
#### Introduction to Representation Learning and Generative Modeling
Tue, March 31st  | 
  * What makes a good representation? How do we know that we found one? 
  * Generative modeling: density estimation, uncertainty modeling 
  * Representation learning: task-relevant encoding 
  * Surrogate tasks: compression, denoising, imputation 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=530467b4-9e25-42d2-a271-b3df01003244)
  * [Slides ](https://drive.google.com/file/d/191YmPcB0_SD-uYs8i3ofgEy4bIj6yfPM/view?usp=drive_link)


  * pset 5 out

 |  
| 
#### Diffusion 1
Thu, April 2nd  | 
  * Mathematical Foundations of Diffusion Models 
  * ODE and SDE perspective of Diffusion 
  * Score Matching and Flow 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=491eee66-6b35-4c38-ae9f-b3df01003261)
  * [Slides ](https://drive.google.com/file/d/1gPtS2QUyawx4gKt0AczW04tH2p8toBbc/view?usp=drive_link)

 |  
| 
#### Diffusion 2
Tue, April 7th  | 
  * Classifier-Free Guidance 
  * Case study: SOTA models in image and video generation 
  * SOTA architectures 

 | 
  * [Recording ](https://mit.hosted.panopto.com/Panopto/Pages/Viewer.aspx?id=9712d05d-538f-40dd-97eb-b3df0100327a)
  * [Slides ](https://drive.google.com/file/d/1dlQGT0YF2fcROIjPfdkd9wp1NoLOPrlW/view?usp=drive_link)

 |  
| 
#### Midterm quiz
Thu, April 9th quiz  
 | 
  * In-class pen-and-paper, closed-book quiz. 

 |   |  
| 
#### Diffusion Models 3
Tue, April 14th  | 
  * A spectral perspective on image and video diffusion 
  * Why do Diffusion Models generalize? 

 | 
  * [Slides ](https://drive.google.com/file/d/1F3Z3mdEO5_7lStNPYyYodBGN5jGfXEY0/view?usp=drive_link)


  * pset 5 due

  

  * Project proposal due

 |  
| 
#### Sequence Generative Models
Thu, April 16th  | 
  * Auto-regressive and full-sequence models 
  * Compounding errors and stability 
  * Diffusion Forcing 
  * History Guidance 

 |   |  
| 
#### Sequence Generative Models II
Tue, April 21st  | 
  * Another perspective on Sequence generation 
  * History Guidance 

 |   |  
| 
#### Non-Generative Representation Learning (Self-supervised learning)
Thu, April 23rd  | 
  * Alternative representation learning techniques 
  * Applications in computer vision 

 |   |  
| 
#### Guest lecture: Mathilde Cornet
Tue, April 28th  |   |   |  
| 
###  Module 3: Module 3: Vision for Embodied Agents
 |  
| 
#### Introduction to Robotic Perception
Thu, April 30th  | 
  * Definition and challenges of embodied agents 
  * Intersection with vision 
  * Controlling Robots from Vision 

 |   |  
| 
#### TBD
Tue, May 5th  |   |   |  
| 
#### Learning Skills from Demonstrations
Thu, May 7th  | 
  * Behavior Cloning and Imitation Learning from Vision 

 |   |  
| 
#### TBD
Tue, May 12th  |   | 
  * Final project due

 |  
### Related Courses and Credits
  * [ **Computer Vision**](https://www.cs.ox.ac.uk/teaching/courses/2024-2025/vision/)  
Oxford University, Prof. Christian Rupprecht 
  * [ **FFTs in Graphics and Vision**](https://www.cs.jhu.edu/~misha/Spring23/)  
Johns Hopkins University, Prof. Misha Kazhdan 
  * [ **Learning for 3D Vision**](https://learning3d.github.io/)  
CMU 16-889, Prof. Shubham Tulsiani 
  * [ **Advances in Computer Vision**](http://6.869.csail.mit.edu/sp22/)  
MIT 6.819/6.869, Profs. Bill Freeman, Phillip Isola, Antonio Torralba 
  * [ **Deep Learning II, Part "Geometric Deep Learning"**](https://uvadl2c.github.io/)  
University of Amsterdam 52042DEL6Y, Prof. Erik Bekkers 
  * [ **Computer Graphics in the Era of AI**](http://cs348i.stanford.edu/)  
Stanford CS348I, Profs. C. Karen Liu and Jiajun Wu 
  * [ **Computer Vision**](https://uni-tuebingen.de/fakultaeten/mathematisch-naturwissenschaftliche-fakultaet/fachbereiche/informatik/lehrstuehle/autonomous-vision/lectures/computer-vision/)  
University of Tübingen ML-4360, Prof. Andreas Geiger 


### Image Attribution
The header background image is by Los Angeles–based designer and photographer [Anastasiya Badun](https://bio.site/_badun), shared [unsplash](https://unsplash.com/photos/a-close-up-of-an-eye-with-a-black-background-2wDxCCw83HM).
We chose the image because of its thematic link to vision, and its overall abstract aesthetic.
© 2021 – 2026 Scene Representation Group
[About MIT](https://web.mit.edu/ "Learn more about MIT") [ About CSAIL ](https://www.csail.mit.edu/ "Learn more about CSAIL")
