# 6.S058 Course Materials

https://introtocv.github.io/materials.html

| [![](https://introtocv.github.io/images/cv_logo1.jpg)](http://groups.csail.mit.edu/vision)  |   
### MIT CSAIL
6.S058: Introduction to Computer Vision  | [![](https://introtocv.github.io/images/cv_logo2.jpg)](http://groups.csail.mit.edu/vision)  |  
| --- | --- | --- |  
| 
### Spring 2026
 |  
| [[Home](https://introtocv.github.io/index.html) | [Policy](https://introtocv.github.io/policy.html) | [Schedule](https://introtocv.github.io/schedule.html) | [Course Materials](https://introtocv.github.io/materials.html) | [Final Project](https://introtocv.github.io/project.html) | [Piazza ](https://piazza.com/class/mkwwzxso70q31t/) | [Canvas](https://canvas.mit.edu/courses/37847) ]  |  
###  Books
Text Book:
  * **[TIF]** Torralba, Isola, Freeman, _[Foundations of Computer Vision](https://mitpress.mit.edu/9780262048972/foundations-of-computer-vision/), _MIT Press, 2024 ([online version](https://visionbook.mit.edu/))


Computer vision:
  * **[Sz]** Szeliski, _[Computer Vision: Algorithms and Applications](http://www.amazon.com/Computer-Vision-Algorithms-Applications-Science/dp/1848829345/), _Springer, 2010 ([online draft](http://szeliski.org/Book/))
  * **[HZ]** Hartley and Zisserman, _[Multiple View Geometry in Computer Vision](http://www.amazon.com/Multiple-View-Geometry-Computer-Vision/dp/0521540518)_ , Cambridge University Press, 2004
  * **[FP]** Forsyth and Ponce, _[Computer Vision: A Modern Approach](http://www.amazon.com/Computer-Vision-Approach-David-Forsyth/dp/0130851981)_ , Prentice Hall, 2002
  * **[Pa]** Palmer, [Vision Science](http://www.amazon.com/Vision-Science-Phenomenology-Stephen-Palmer/dp/0262161834/), MIT Press, 1999


Learning:
  * **[GBC]** Goodfellow, Bengio, Courville, _[Deep Learning](https://www.deeplearningbook.org/), _MIT Press, 2016
  * **[Mi]** Mitchel, _[Machine Learning](http://www.amazon.com/Machine-Learning-Tom-M-Mitchell/dp/0070428077/), _McGraw-Hill, 1997
  * **[DHS]** Duda, Hart and Stork, _[Pattern Classification (2nd Edition)](http://www.amazon.com/Pattern-Classification-2nd-Richard-Duda/dp/0471056693)_ , Wiley-Interscience, 2000
  * **[SB]** Sutton & Barto, _[On-line book. The classic reference to the field of reinforcement learning.](http://incompleteideas.net/book/the-book-2nd.html)_


Graphical models:
  * **[KF]** Koller and Friedman, _[Probabilistic Graphical Models: Principles and Techniques](http://www.amazon.com/Probabilistic-Graphical-Models-Principles-Computation/dp/0262013193/)_ , MIT Press, 2009


###  Resources
Image datasets:
  * [Labelme](http://labelme.csail.mit.edu/): an online annotation tool to build image databases for computer vision research
  * [OpenSurfaces](http://opensurfaces.cs.cornell.edu/): a large database of annotated surfaces created from real-world consumer photographs.
  * [ImageNet](http://http://image-net.org/): a large-scale image dataset for visual recognition organized by [WordNet](http://wordnet.princeton.edu/) hierarchy
  * [ADE20K Dataset](https://groups.csail.mit.edu/vision/datasets/ADE20K/): a benchmark for scene and instance segmentation, with pixelwise semantic annotations
  * [Places Database](http://places.csail.mit.edu/): a scene-centric database with 205 scene categories and 2.5 millions of labelled images
  * [NYU Depth Dataset v2](http://cs.nyu.edu/~silberman/datasets/nyu_depth_v2.html): a RGB-D dataset of segmented indoor scenes
  * [Microsoft COCO](http://mscoco.org/): a new benchmark for image recognition, segmentation and captioning
  * [Flickr100M](http://yahoolabs.tumblr.com/post/89783581601/one-hundred-million-creative-commons-flickr-images): 100 million creative commons Flickr images
  * [Labeled Faces in the Wild](http://vis-www.cs.umass.edu/lfw/): a dataset of 13,000 labeled face photographs
  * [Human Pose Dataset](http://human-pose.mpi-inf.mpg.de/): a benchmark for articulated human pose estimation
  * [YouTube Faces DB](http://www.cs.tau.ac.il/~wolf/ytfaces/): a face video dataset for unconstrained face recognition in videos
  * [UCF101](http://crcv.ucf.edu/data/UCF101.php): an action recognition data set of realistic action videos with 101 action categories
  * [HMDB-51](http://serre-lab.clps.brown.edu/resource/hmdb-a-large-human-motion-database/): a large human motion dataset of 51 action classes


Top computer vision conferences and papers:
  * [CVPR](http://www.pamitc.org/cvpr15/accepted_papers.php): IEEE Conference on Computer Vision and Pattern Recognition
  * [ICCV](http://www.cvpapers.com/iccv2013.html): International Conference on Computer Vision
  * [ECCV](http://eccv2014.org/accepted-papers/): European Conference on Computer Vision
  * [NeurIPS](http://nips.cc/Conferences/2014/Program/accepted-papers.php): Neural Information Processing Systems


Related courses:
  * [Introduction to Computer Vision](http://www.cs.brown.edu/courses/cs143/), by Michael Black
  * [Learning-Based Methods in Vision](http://www.cs.cmu.edu/%7Eefros/courses/LBMV07/), by Alyosha Efros
  * [Computer Vision](http://www.cs.utexas.edu/%7Egrauman/courses/fall2009/main.htm), by Kristen Grauman
  * [Computer Vision](http://www.cs.nyu.edu/%7Efergus/), by Rob Fergus
  * [Introduction to Computer Vision](http://vision.stanford.edu/teaching/cs223b/), by Fei-Fei Li


Other resources:
  * [The Computer Vision Industry](http://www.cs.ubc.ca/spider/lowe/vision.html)
  * [MatCovNet](http://www.vlfeat.org/matconvnet/)-CNN Toolbox in Matlab
  * [Pytorch](https://pytorch.org/)
  * [Tensorflow](https://www.tensorflow.org/)


