# Stanford CS 224N | Project Reports

https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/project.html

[CS224N Home](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/index.html)
  * [Coursework](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/index.html#coursework)
  * [Schedule](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/index.html#schedule)
  * [Office Hours](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/office_hours.html)
  * [Final projects](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/project.html)
  * [Lecture Videos](https://canvas.stanford.edu/courses/164570/external_tools/3367)
  * [Ed Forum](https://edstem.org/us/courses/51053)


[ ![](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/images/stanford-nlp-logo-new.jpg) ](http://nlp.stanford.edu/) [ ![](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/images/stanfordlogo.jpg) ](http://stanford.edu/)
# CS224N: Natural Language Processing with Deep Learning
### Stanford / Winter 2024
## Final poster session
We thank our sponsor, Sky9 Capital, for supporting the poster session!   
The poster session was held at the AOERC basketball courts from 7-10 PM on March 18th, 2024. 
![](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/images/sky9.png)
## Outstanding Projects
### Student choice for best poster
  * **[The Beat Goes On: Symbolic Music Generation with Text Controls](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JavokhirBArifovNathanaelJamesCadicamoPhilipAndrewBaillargeon.pdf)**. Javokhir B Arifov, Nathanael James Cadicamo, Philip Andrew Baillargeon 
  * **[Identifying and Neutralizing Gender Bias from Text](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DanteEmanuelDanelianLiamMichaelSmithMayaBedge.pdf)**. Dante Emanuel Danelian, Liam Michael Smith, Maya Bedge 
  * **[Lyricade: An Integrated Acoustic Signal-Processing Transformer for Lyric Generation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArnavSomayajiKrishnamoorthiKlaraBjorkAndraThomasRheaMalhotra.pdf)**. Arnav Somayaji Krishnamoorthi, Klara Bjork Andra-Thomas, Rhea Malhotra 


### Outstanding custom project reports
  * **[An Inside Look Into How LLMs Fail to Express Complete Certainty: Are LLMs Purposely Lying?](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JoshuaDeGuzmanFajardo.pdf)** Joshua De Guzman Fajardo 
  * **[Count Your Words Before They Hatch: Investigating Word Count Control](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KatherineLi.pdf)**. Katherine Li 
  * **[Identifying and Neutralizing Gender Bias from Text](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DanteEmanuelDanelianLiamMichaelSmithMayaBedge.pdf)**. Dante Emanuel Danelian, Liam Michael Smith, Maya Bedge 
  * **[Self-Improvement for Math Problem-Solving in Small Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArtyomShaposhnikovRobertoGarciaTorresShubhraMishra.pdf)**. Artyom Shaposhnikov, Roberto Garcia Torres, Shubhra Mishra 
  * **[Increasing the Efficiency of the Sophia Optimizer: Continuous Adaptive Information Averaging](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CaiaMaiCostelloJasonDanielLazar.pdf)**. Caia Mai Costello, Jason Daniel Lazar 
  * **[SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexMuzioAlexSunChuranHe.pdf)**. Alex Muzio, Alex Sun, Churan He 
  * **[Claim-level Uncertainty Estimation through Graph](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MingjianJiang.pdf)**. Mingjian Jiang 
  * **[CapNet: Making Science More Accessible via a Neural Caption Generator](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AmanLadiaTirthDharmeshSurti.pdf)**. Aman Ladia, Tirth Dharmesh Surti 
  * **[20 Questions: Efficient Adaptation for Individualized LLM Personalization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MichaelJosephRyan.pdf)**. Michael Joseph Ryan 
  * **[Improving Low-Resource POS Tagging with Transfer Learning: A Case in Cantonese](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ManhDDaoTingLinTrevorWilliamCarrell.pdf)**. Manh D Dao, Ting Lin, Trevor William Carrell 
  * **Probing to Interpret the Mysterious Phenomenon of In-Context Learning in LLMs**. Junyi Tao 
  * **[Fine-tuning CodeLlama-7B on Synthetic Training Data for Fortran Code Generation using PEFT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AndrewCShiSohamGovandeTaeukKang.pdf)**. Andrew C Shi, Soham Govande, Taeuk Kang 
  * **[Clinically relevant summarization of multimodal emergency medical data](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElsaBismuthJanMichaelKrauseLucasALeanza.pdf)**. Elsa Bismuth, Jan Michael Krause, Lucas A Leanza 
  * **[Agent Retrieval on Textual and Relational Knowledge Bases](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ShiyuZhao.pdf)**. Shiyu Zhao 
  * **Searching your Backpack: Information Retrieval with Backpack Language Models**. Brian Christopher Johnson 
  * **[Automated Extraction of ICD-10 Diagnosis Codes from Clinical Notes](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArjunJainDevanshuLadsariaRishiRajVerma.pdf)**. Arjun Jain, Devanshu Ladsaria, Rishi Raj Verma 
  * **[Prototype-then-Refine: A Neurosymbolic Approach for Improved Logical Reasoning with LLMs](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BassemAkoushHashemElezabi.pdf)**. Bassem Akoush, Hashem Elezabi 
  * **[Enhancing Factuality in Language Models through Knowledge-Guided Decoding](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JirayuBurapacheep.pdf)**. Jirayu Burapacheep 
  * **[On Fairness Implications and Evaluations of Low-Rank Adaptation of Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ZhoujieDing.pdf)**. Zhoujie Ding 
  * **[Towards Natural Language Reasoning for Unified Robotics Description Format Files](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AakashMishraAustinAnilPatelNeilNie.pdf)**. Aakash Mishra, Austin Anil Patel, Neil Nie 
  * **[AI-Driven Fashion Cataloging: Transforming Images into Textual Descriptions](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NishantGopinathSiYiMa.pdf)**. Nishant Gopinath, Si Yi Ma 
  * **[Model Mixture: Merging Task-Specific Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElizabethZhuSherryXie.pdf)**. Elizabeth Zhu, Sherry Xie 


### Outstanding default project reports
  * **[Methods to Improve Downstream Generalization of minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RamgopalVenkateswaran.pdf)**. Ramgopal Venkateswaran 
  * **[Multi-BERT: A Multi-Task BERT Approach with the Variation of Projected Attention Layer](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaijingZhang.pdf)**. Haijing Zhang 
  * **[BERT and Beyond: A Study of Multitask Learning Strategies for NLP](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/FebieJaneLinJackPLe.pdf)**. Febie Jane Lin, Jack P Le 
  * **[Enhanced TreeBERT: High-Performance, Computationally Efficient Multi-Task Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/PannSripitakThanawanAtchariyachanvanit.pdf)**. Pann Sripitak, Thanawan Atchariyachanvanit 
  * **[Beyond Fine-tuning: Iterative Ensemble Strategies for Enhanced BERT Generalizability](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MeganDassRiyaDulepetShreyaDSouza.pdf)**. Megan Dass, Riya Dulepet, Shreya D'Souza 
  * **[Semantic Symphonies: BERTrilogy and BERTriad Ensembles](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaoyiDuanYaohuiZhang.pdf)**. Haoyi Duan, Yaohui Zhang 
  * **[Good Things Come to Those Who Weight: Effective Pairing Strategies for Multi-Task Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NachatJatusripitakPawanWirawarn.pdf)**. Nachat Jatusripitak, Pawan Wirawarn 
  * **[Jack of All Trades, Master of Some: Improving BERT for Multitask Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChijiokeMgbahurikeIddahMlauziKwameOcran.pdf)**. Chijioke Mgbahurike, Iddah Mlauzi, Kwame Ocran 
  * **[SMART loss vs DeBERTa](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MichaelLiuMichaelPhillipHayashiRobertoLobatoLopez.pdf)**. Michael Liu, Michael Phillip Hayashi, Roberto Lobato Lopez 
  * **[QuarBERT: Optimizing BERT with Multitask Learning and Quartet Ensemble](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CarrieGuErickaLiuZixinLi.pdf)**. Carrie Gu, Ericka Liu, Zixin Li 
  * **[MT-DNN with SMART Regularisation and Task-Specific Head to Capture the Pairwise and Contextually Significant Words Interplay](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaoyuWang.pdf)**. Haoyu Wang 
  * **[Few-Shot Prompt-Tuning: An Extension to a Finetuning Alternative](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JamesJMoriceSamuelEdwardKwok.pdf)**. James J Morice, Samuel Edward Kwok 
  * **[AdaptBert: Parameter Efficient Multitask Bert](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JietingQiuShwetaAgrawal.pdf)**. Jieting Qiu, Shweta Agrawal 


## Custom Projects  
| [Fast, Interpretable AI-Generated Text Detection Using Style Embeddings](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GiuliaZoeSocolofRitikaKacholia.pdf)  | Giulia Zoe Socolof, Ritika Kacholia  |  
| --- | --- |  
| [CapNet: Making Science More Accessible via a Neural Caption Generator](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AmanLadiaTirthDharmeshSurti.pdf)  | Aman Ladia, Tirth Dharmesh Surti  |  
| [Do LLMs exhibit Nominal Compound Understanding, or just Nominal Understanding?](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElijahSongNathanAndrewChi.pdf)  | Elijah Song, Nathan Andrew Chi  |  
| [Using Language Model for Emission Factor Mapping](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AmarkumarKallappaGadkariEstebanJoseBarreroHernandezGeraldChavinKang.pdf)  | Amarkumar Kallappa Gadkari, Esteban Jose Barrero-Hernandez, Gerald Chavin Kang  |  
| [Affective Emotional Layer for Conversational LLM Agents](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AdityaBoraNikhilSuresh.pdf)  | Aditya Bora, Nikhil Suresh  |  
| [Synthesized Strategy for Mental Health Support](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DavidYuanEvelynSong.pdf)  | David Yuan, Evelyn Song  |  
| [Text2Gloss: Translation into Sign Language Gloss with Transformers](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JennaSaraMansuetoLukeCBabbitt.pdf)  | Jenna Sara Mansueto, Luke C Babbitt  |  
| [BondBERT: An ensemble-based model for named entity recognition in materials science texts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BellaCrouch.pdf)  | Bella Crouch  |  
| [The Beat Goes On: Symbolic Music Generation with Text Controls](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JavokhirBArifovNathanaelJamesCadicamoPhilipAndrewBaillargeon.pdf)  | Javokhir B Arifov, Nathanael James Cadicamo, Philip Andrew Baillargeon  |  
| [Near-Infinite Sub-Quadratic Convolutional Attention](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JamesPoetzscher.pdf)  | James Poetzscher  |  
| [Llama-UL2: Emerging New Capabilities with Continued Pretraining using UL2](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JasonShuYangWang.pdf)  | Jason Shu-Yang Wang  |  
| [Feedback or Autonomy? Analyzing LLMs’ Ability to Self-Correct](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KaiMicaFronsdal.pdf)  | Kai Mica Fronsdal  |  
| [Brain-to-Text](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EthanTrepka.pdf)  | Ethan Trepka  |  
| [Finetuning Provides a Window Into Transformer Circuits](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SachalSohanSrivastavaMalick.pdf)  | Sachal Sohan Srivastava-Malick  |  
| [Funding Sources and Values of NLP Research](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ParthSarinPatriciaWeiVyomaRaman.pdf)  | Parth Sarin, Patricia Wei, Vyoma Raman  |  
| [Understanding Complex Emotions in Sentences](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HannahRachelLevinJuanPabloTrianaMartinezSamirAgarwala.pdf)  | Hannah Rachel Levin, Juan Pablo Triana Martinez, Samir Agarwala  |  
| [Lyricade: An Integrated Acoustic Signal-Processing Transformer for Lyric Generation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArnavSomayajiKrishnamoorthiKlaraBjorkAndraThomasRheaMalhotra.pdf)  | Arnav Somayaji Krishnamoorthi, Klara Bjork Andra-Thomas, Rhea Malhotra  |  
| [Training a Chinese RapStar: Applying Rapformer Model to Generate Chinese Rap Lyrics](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BihanLiuMikeYang.pdf)  | Bihan Liu, Mike Yang  |  
| [20 Questions: Efficient Adaptation for Individualized LLM Personalization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MichaelJosephRyan.pdf)  | Michael Joseph Ryan  |  
| [LLMs for Google Maps](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SukrutOak.pdf)  | Sukrut Oak  |  
| [BERT one-shot movie recommender system](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChiTrungNguyen.pdf)  | Chi Trung Nguyen  |  
| [Over-Complicating GPT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DanielLiYang.pdf)  | Daniel Li Yang  |  
| [Semantics of Empire: A Neural Machine Translation Approach for Ottoman Turkish Texts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MerveTekgurler.pdf)  | Merve Tekgurler  |  
| [Outrageously Fast LLMs: Faster Inference and Fine-Tuning with Moefication and LoRA](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChiYoTsaiJayMartin.pdf)  | Chi Yo Tsai, Jay Martin  |  
| [Speaking the Language of Sight](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JeanRodmondJuniorLaguerreVickyWu.pdf)  | Jean Rodmond Junior Laguerre, Vicky Wu  |  
| [Reliable Ambient Intelligence Through Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DavidDai.pdf)  | David Dai  |  
| [Predicting Big Brother Brasil 2024 Evictions Through Sentiment Analysis of Tweets](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LauraFiuzaDubugras.pdf)  | Laura Fiuza Dubugras  |  
| [Count Your Words Before They Hatch: Investigating Word Count Control](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KatherineLi.pdf)  | Katherine Li  |  
| [Novelty: Optimizing StreamingLLM for Novel Plot Generation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JoyceChuyiChenMeganMou.pdf)  | Joyce Chuyi Chen, Megan Mou  |  
| [Multimodal MoE for InfographicsVQA](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ManoloAlvarez.pdf)  | Manolo Alvarez  |  
| [Patent Acceptance Prediction With Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MiguelGerenaRiveraakaylahackson.pdf)  | Miguel Gerena Rivera, Akayla Hackson  |  
| [Salience-Based Adversarial Attacks for Empirical Evaluation of NLP Classification Robustness](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/FletcherLeeNewell.pdf)  | Fletcher Lee Newell  |  
| [Computing semantic textual similarity through transformer-based encoders and combining multiple content similarity measures](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EthanYiKoKaysonTakaHansenPengHaoLu.pdf)  | Ethan Yi Ko, Kayson Taka Hansen, Peng Hao Lu  |  
| [AI-Driven Fashion Cataloging: Transforming Images into Textual Descriptions](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NishantGopinathSiYiMa.pdf)  | Nishant Gopinath, Si Yi Ma  |  
| [A Multi-tiered Approach to Debiasing Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AmanKansalSaanviChawlashreyashankar.pdf)  | Aman Kansal, Saanvi Chawla, Shreya Shankar  |  
| [Deciphering the Dynamics of Reddit Comment Popularity](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AbdulwahabOmira.pdf)  | Abdulwahab Omira  |  
| [Natural Language Enhanced Neural Program Synthesis for Abstract Reasoning Task](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/YijiaWang.pdf)  | Yijia Wang  |  
| [Self-Improvement for Math Problem-Solving in Small Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArtyomShaposhnikovRobertoGarciaTorresShubhraMishra.pdf)  | Artyom Shaposhnikov, Roberto Garcia Torres, Shubhra Mishra  |  
| [Guided Image Concept Decomposition using Textual Inversion](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/YvetteYinyinLin.pdf)  | Yvette Yinyin Lin  |  
| [From Beethoven to Beyoncé: A Deep Learning Approach to Music Genre Classification](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DominicJosephDeMarcoEricMartzReginaTHTa.pdf)  | Dominic Joseph DeMarco, Eric Martz, Regina T.H. Ta  |  
| [Identifying and Neutralizing Gender Bias from Text](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DanteEmanuelDanelianLiamMichaelSmithMayaBedge.pdf)  | Dante Emanuel Danelian, Liam Michael Smith, Maya Bedge  |  
| [Efficient Alignment of Medical Language Models using Direct Preference Optimization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BrendanMurphy.pdf)  | Brendan Murphy  |  
| [Boosting Embodied Reasoning in LLMs in Multi-agent Mixed Incentive Environments](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AgamMohanSinghBhatia.pdf)  | Agam Mohan Singh Bhatia  |  
| [Virgilian Poetry Generation with LSTM Networks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AugustWyattBurtonJonathanDanielMerchan.pdf)  | August Wyatt Burton, Jonathan Daniel Merchan  |  
| [Agent Retrieval on Textual and Relational Knowledge Bases](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ShiyuZhao.pdf)  | Shiyu Zhao  |  
| [High-fidelity Human Representation for Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BrianXuHenryJinWeng.pdf)  | Brian Xu, Henry Jin Weng  |  
| [Improving Human-LLM Interactions by Redesigning the Chatbot](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RashonPoole.pdf)  | Rashon Poole  |  
| [GRAPHGEM: Improving Graph Reasoning in Language Models with Synthetic Data](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/TonySun.pdf)  | Tony Sun  |  
| [Mathematical Reasoning Through LLM Finetuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KarthikVinaySeetharamanYashMehta.pdf)  | Karthik Vinay Seetharaman, Yash Mehta  |  
| [Effects of Pre-training and Fine-tuning Time on the Linear Connectivity of Language Models for Natural Language Inference](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KushalThaman.pdf)  | Kushal Thaman  |  
| [Cross-Lingual Summarization of Notice to Air Missions (NOTAMs)](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ZixiLiu.pdf)  | Zixi Liu  |  
| [GILgaMeSH: Glyph-Interpreting Language Models for Sumerian History](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ColeSimmons.pdf)  | Cole Simmons  |  
| [A Good Novelist Should be a Good Coder: From Language Critics to Automatic Code Generation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BrianManuelMunozMushengLin.pdf)  | Brian Manuel Munoz, Mu-sheng Lin  |  
| [Wrestling Mamba: Exploring Early Fine-Tuning Dynamics on Mamba and Transformer Architectures](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DanielGuoLucasEmmanuelBrennanAlmaraz.pdf)  | Daniel Guo, Lucas Emmanuel Brennan-Almaraz  |  
| [An Inside Look Into How LLMs Fail to Express Complete Certainty: Are LLMs Purposely Lying?](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JoshuaDeGuzmanFajardo.pdf)  | Joshua De Guzman Fajardo  |  
| [Spatial-Enhanced Summarization of Placement Preferences For Robot-Action Personalization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SidPotti.pdf)  | Sid Potti  |  
| [RoSA Text Style Transfer & Evaluation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArnavGuptaAyaanNaveedMalikMacVincentSomtochukwuAghaOko.pdf)  | Arnav Gupta, Ayaan Naveed Malik, MacVincent Somtochukwu Agha-Oko  |  
| [SUPaHOT: Universally Scalable and Private Method to Demystify FHIR Health Records](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HamedHekmatMichaelNBrockmanNinaBoord.pdf)  | Hamed Hekmat, Michael N Brockman, Nina Boord  |  
| [The Development of Facticity—from Preliminary Findings to Accepted Implicit Knowledge: Case Studies](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JingruoSunTianyuDuYuzeSui.pdf)  | Jingruo Sun, Tianyu Du, Yuze Sui  |  
| [Short Text Classification of Political Reddit Posts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ShirleyCheng.pdf)  | Shirley Cheng  |  
| [Predicting Patent Litigation Risk Using RoBERTa and Metadata Augmentation Techniques](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BrianParkNikitaBhardwajSimoneYiYiHsu.pdf)  | Brian Park, Nikita Bhardwaj, Simone Yi-Yi Hsu  |  
| [Puzzle in a Haystack: Understanding & Enhancing Long Context Reasoning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JessicaChudnovskySalmanAbdullahSudharsanSundar.pdf)  | Jessica Chudnovsky, Salman Abdullah, Sudharsan Sundar  |  
| [Graph-based Logical Reasoning for Legal Judgement](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EinJun.pdf)  | Ein Jun  |  
| [Enabling Cross-Linguistic Compatibility in Image Generation: Text Embedding Alignment Techniques for CLIP Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BofeiZhu.pdf)  | Bofei Zhu  |  
| [Autoformalization with Backtranslation: Training an Automated Mathematician](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JakobNordhagen.pdf)  | Jakob Nordhagen  |  
| [Tree-Based Retrieval Using Gaussian Statistics](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IrfanNafiLukeJongsungParkRayHotate.pdf)  | Irfan Nafi, Luke Jongsung Park, Ray Hotate  |  
| [Predicting Protein-Protein Interaction via Protein Textual Description using Large Language Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KhoaHoang.pdf)  | Khoa Hoang  |  
| [Information Dense Question Answering for RLHF](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChetAnandBhateja.pdf)  | Chet Anand Bhateja  |  
| [Evaluating the Culture-awareness in Pre-trained Language Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RyanLiYutongZhangZhiyuXie.pdf)  | Ryan Li, Yutong Zhang, Zhiyu Xie  |  
| [Stanford CS Course + Quarter Classification based on CARTA Reviews](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EnokChoeJubenRana.pdf)  | Enok Choe, Juben Rana  |  
| [Robotics Tasks Generation through Factorization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AngelaYiRanajitGangopadhyayYihanZhou.pdf)  | Angela Yi, Ranajit Gangopadhyay, Yihan Zhou  |  
| [Learning Strategic Play with Language Agents in Text-Adventure Games](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MirandaLinLiNicBecker.pdf)  | Miranda Lin Li, Nic Becker  |  
| [Detect failure root cause and predict faults from software logs](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SandipPal.pdf)  | Sandip Pal  |  
| [Direct Clinician Preference Optimization: Clinical Text Summarization via Expert Feedback-Integrated LLMs](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MikeTimmermanOnatDalmazTimNiklausReinhart.pdf)  | Mike Timmerman, Onat Dalmaz, Tim Niklaus Reinhart  |  
| [Exploring Unsupervised Machine Translation for Highly Under resourced Languages(Hausa)](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SajidOmarFarookZouberouSayibou.pdf)  | Sajid Omar Farook, Zouberou Sayibou  |  
| [Automated Extraction of ICD-10 Diagnosis Codes from Clinical Notes](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArjunJainDevanshuLadsariaRishiRajVerma.pdf)  | Arjun Jain, Devanshu Ladsaria, Rishi Raj Verma  |  
| [Detecting Misinformation in News Articles via Natural Language Processing](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SiyaGoelThuLeTiaVasudeva.pdf)  | Siya Goel, Thu Le, Tia Vasudeva  |  
| [Model Mixture: Merging Task-Specific Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElizabethZhuSherryXie.pdf)  | Elizabeth Zhu, Sherry Xie  |  
| [Prototype-then-Refine: A Neurosymbolic Approach for Improved Logical Reasoning with LLMs](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BassemAkoushHashemElezabi.pdf)  | Bassem Akoush, Hashem Elezabi  |  
| [Minimal Clues for Maximal Understanding: Solving Linguistic Puzzles with RNNs, Transformers, and LLMs](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnaviBaddepudiEmmaWangIshanKhare.pdf)  | Anavi Baddepudi, Emma Wang, Ishan Khare  |  
| [Difficulty-Controllable Text Generation Aligned to Human Preferences](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AllisonGumanNikhilPanditUdayanMandal.pdf)  | Allison Guman, Nikhil Pandit, Udayan Mandal  |  
| [Optimized Linear Attention for TPU Hardware](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GabraelLevine.pdf)  | Gabrael Levine  |  
| [Claim-level Uncertainty Estimation through Graph](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MingjianJiang.pdf)  | Mingjian Jiang  |  
| [Extracting Material Measurement Knowledge Graphs from Academic Research Papers](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArthurCerqueiraCampello.pdf)  | Arthur Cerqueira Campello  |  
| [GaS – Graph and Sequence Modeling for Web Agent Pathing](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArantxaRamosdelValleMarcelArzhangHeshmatiRoedMatthewNoto.pdf)  | Arantxa Ramos del Valle, Marcel Arzhang Heshmati Roed, Matthew Noto  |  
| [SpamResponder: Automatic Response System for Voice Phishing](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChaeYoungLee.pdf)  | Chae Young Lee  |  
| [Improving Low-Resource POS Tagging with Transfer Learning: A Case in Cantonese](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ManhDDaoTingLinTrevorWilliamCarrell.pdf)  | Manh D Dao, Ting Lin, Trevor William Carrell  |  
| [Fine-tuning CodeLlama-7B on Synthetic Training Data for Fortran Code Generation using PEFT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AndrewCShiSohamGovandeTaeukKang.pdf)  | Andrew C Shi, Soham Govande, Taeuk Kang  |  
| [Logic-LangChain: Translating Natural Language to First Order Logic for Logical Fallacy Detection](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AbhinavLalwaniIshikaaLunawat.pdf)  | Abhinav Lalwani, Ishikaa Lunawat  |  
| [Comparative Analysis of Preference-Informed Alignment Techniques for Language Model Alignment](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SooWeiKoh.pdf)  | Soo Wei Koh  |  
| [How you can convince ChatGPT the world is flat](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JulianCheng.pdf)  | Julian Cheng  |  
| [Exploring Machine Unlearning in Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JayGuptaLawrenceYChai.pdf)  | Jay Gupta, Lawrence Y Chai  |  
| [MediGANdist: Improving Smaller Model’s Medical Reasoning via GAN-inspired Distillation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AllenKiriroathChauAryanSiddiqui.pdf)  | Allen Kiriroath Chau, Aryan Siddiqui  |  
| [PromptCom: Cost Optimization of Language Models Based on Prompt Complexity](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MengzeGaoYonatanUrman.pdf)  | Mengze Gao, Yonatan Urman  |  
| [Extractive Question Answering On Large Structural Engineering Documents](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AdamUsmaniBangaThomasSounack.pdf)  | Adam Usmani Banga, Thomas Sounack  |  
| [Task-Agnostic Low-Rank Dialectal Adapters for Speech-Text Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AmyShilinGuanAzureSiyiZhouClaireLynnShao.pdf)  | Amy Shilin Guan, Azure Siyi Zhou, Claire Lynn Shao  |  
| [NatuRel: Advancing Relational Understanding in Vision-Language Models with Natural Language Variations](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AshnaKhetanIsabelPazReyesSiehLayaBalajiIyer.pdf)  | Ashna Khetan, Isabel Paz Reyes Sieh, Laya Balaji Iyer  |  
| [More Effectively Searching Trees of Thought for Increased Reasoning Ability in Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KamyarJohnSalahiPranavGurusankarSathyaEdamadaka.pdf)  | Kamyar John Salahi, Pranav Gurusankar, Sathya Edamadaka  |  
| [Improving performance in large language models through diversity of thoughts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CorneliaWeinzierlSreethuSuraSugunaVarshiniVelury.pdf)  | Cornelia Weinzierl, Sreethu Sura, Suguna Varshini Velury  |  
| [Increasing the Efficiency of the Sophia Optimizer: Continuous Adaptive Information Averaging](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CaiaMaiCostelloJasonDanielLazar.pdf)  | Caia Mai Costello, Jason Daniel Lazar  |  
| [Unstructured Data Abstraction utilizing Selective Prediction-Oriented Neural Networks in Healthcare Settings](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EmilyYiduoChenNathanNMohitNicoleTong.pdf)  | Emily Yiduo Chen, Nathan N. Mohit, Nicole Tong  |  
| [Claim Verification for Fictional Narratives with Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JenniferJingXuLaurenYumiKong.pdf)  | Jennifer Jing Xu, Lauren Yumi Kong  |  
| [Enhancing Factuality in Language Models through Knowledge-Guided Decoding](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JirayuBurapacheep.pdf)  | Jirayu Burapacheep  |  
| [Temporal Grounding of Activities using Multimodal Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/YoungCholSong.pdf)  | Young Chol Song  |  
| [AI Lie Detection: Is the Hype Justified?](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JackRyan.pdf)  | Jack Ryan  |  
| [A Comparative Study of Deep Learning Architectures for Long Text Classification in Mental Health](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IvySunSiqiMaYiranFan.pdf)  | Ivy Sun, Siqi Ma, Yiran Fan  |  
| [Towards Natural Language Reasoning for Unified Robotics Description Format Files](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AakashMishraAustinAnilPatelNeilNie.pdf)  | Aakash Mishra, Austin Anil Patel, Neil Nie  |  
| [SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexMuzioAlexSunChuranHe.pdf)  | Alex Muzio, Alex Sun, Churan He  |  
| [Compression Ratio Controlled Text Summarization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ZhengWang.pdf)  | Zheng Wang  |  
| [Multi-Agent Frameworks in Domain-Specific Question Answering Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EthanDuncanHeLiHellmanMariaAngelikaNikitaSpencerLouisPaul.pdf)  | Ethan Duncan He-Li Hellman, Maria Angelika-Nikita, Spencer Louis Paul  |  
| [Understanding Visual Shortcomings of Multimodal Large Language Model Through Training Data Distribution](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlanLiBinxuLi.pdf)  | Alan Li, Binxu Li  |  
| [Clinically relevant summarization of multimodal emergency medical data](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElsaBismuthJanMichaelKrauseLucasALeanza.pdf)  | Elsa Bismuth, Jan Michael Krause, Lucas A Leanza  |  
| [LLMs with Low-Resource Translation: Syriac-to-English Case Study](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AndrewTinLokLee.pdf)  | Andrew Tin-Lok Lee  |  
| [Llama2.pi: Running LLMs on the Bleeding Edge](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MatthewDing.pdf)  | Matthew Ding  |  
| [“Not All Information is Created Equal”: Leveraging Metadata for Enhanced Knowledge Curation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HongMengYamYuchengJiang.pdf)  | Hong Meng Yam, Yucheng Jiang  |  
| [GeoPolitical Risk Predictor](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LucasPBosmanWilliamTobyDenton.pdf)  | Lucas P Bosman, William Toby Denton  |  
| [Multimodal Social Media Sentiment Analysis](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MubarakAliSeyedIbrahimPratyushMuthukumar.pdf)  | Mubarak Ali Seyed Ibrahim, Pratyush Muthukumar  |  
| [Few-Shot Prompt-Tuning: An Extension to a Finetuning Alternative](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JamesJMoriceSamuelEdwardKwok.pdf)  | James J Morice, Samuel Edward Kwok  |  
| [On Fairness Implications and Evaluations of Low-Rank Adaptation of Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ZhoujieDing.pdf)  | Zhoujie Ding  |  
| [Predicting Yelp Star Ratings: An Analysis of Different Models and Fine-Tuned RoBERTa Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RishiAlluriUpamanyuDassVattam.pdf)  | Rishi Alluri, Upamanyu Dass-Vattam  |  
## Default Projects  
| [SMARTer BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexeyAlexandrovichTuzikovNaijingGuoTatianaVeremeenko.pdf)  | Alexey Alexandrovich Tuzikov, Naijing Guo, Tatiana Veremeenko  |  
| --- | --- |  
| [Using Stochastic Layer Dropping as a Regularization Tool to Improve Downstream Prediction Accuracy](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KarthikJetty.pdf)  | Karthik Jetty  |  
| [AdaptBert: Parameter Efficient Multitask Bert](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JietingQiuShwetaAgrawal.pdf)  | Jieting Qiu, Shweta Agrawal  |  
| [An Exploration of Fine-Tuning Techniques on minBERT Optimizations](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GabrielaCortesIrisTFuVictoriaHsieh.pdf)  | Gabriela Cortes, Iris T Fu, Victoria Hsieh  |  
| [Three Heads are Better than One: Implementing Multiple Models with Task-Specific BERT Heads](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MattAlexanderKaplanPreritChoudharySinaMohammadi.pdf)  | Matt Alexander Kaplan, Prerit Choudhary, Sina Mohammadi  |  
| [Multitask BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BradleyHuShannonXiao.pdf)  | Bradley Hu, Shannon Xiao  |  
| [2-Tier SimCSE: Elevating BERT for Robust Sentence Embeddings](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AubreyWangCandiceWangZiranZhou.pdf)  | Aubrey Wang, Candice Wang, Ziran Zhou  |  
| [Sentence-BERT-inspired Improvements to minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RajVPabari.pdf)  | Raj V Pabari  |  
| [minBERT and Downstream Tasks Optimization with Disentangled Attention](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JeremyLinfieldSeanBai.pdf)  | Jeremy Linfield, Sean Bai  |  
| [Choose Your PALs Wisely](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ZachPeterRotzal.pdf)  | Zach Peter Rotzal  |  
| [Good Things Come to Those Who Weight: Effective Pairing Strategies for Multi-Task Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NachatJatusripitakPawanWirawarn.pdf)  | Nachat Jatusripitak, Pawan Wirawarn  |  
| [Semantic Symphonies: BERTrilogy and BERTriad Ensembles](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaoyiDuanYaohuiZhang.pdf)  | Haoyi Duan, Yaohui Zhang  |  
| [MinBERT and PALs: Multi-Task Leaning for Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/TetsuyaHayashi.pdf)  | Tetsuya Hayashi  |  
| [minBERT Multi Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AugustinBoissierMaximePedron.pdf)  | Augustin Boissier, Maxime Pedron  |  
| [Simple Contrastive Learning for Multitask Finetuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnnieZZhuGuiDavidKhaingSuMon.pdf)  | Annie Z Zhu, Gui David, Khaing Su Mon  |  
| [minBERT, NLP Tasks, and More](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/TiankaiYan.pdf)  | Tiankai Yan  |  
| [Exploring Pretraining, Finetuning and Regularization for Multitask Learning of minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/WeichengSongXinyuHuZhiyinPan.pdf)  | Weicheng Song, Xinyu Hu, Zhiyin Pan  |  
| [Enhancing BERT for NLP Tasks: Pretraining, Fine-tuning, and Model Augmentation](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BryantPerkinsDylanRyanDipasupil.pdf)  | Bryant Perkins, Dylan Ryan Dipasupil  |  
| [Loss Weighting in Multi-Task Language Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnnaLittle.pdf)  | Anna Little  |  
| [Optimizing minBERT on Downstream Tasks Using Pretraining and Siamese Network Architecture](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EdwinAntonioPua.pdf)  | Edwin Antonio Pua  |  
| [Learning by Prediction and Diversity with BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexLin.pdf)  | Alex Lin  |  
| [Exploring Challenges in Multi-task BERT Optimization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IsabelMichel.pdf)  | Isabel Michel  |  
| [SMARTCS: Additional Pretraining and Robust Finetuning on BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AyeshaKhawajaRachelSinaiClintonYasmineFatimaMabene.pdf)  | Ayesha Khawaja, Rachel Sinai Clinton, Yasmine Fatima Mabene  |  
| [Even Language Models Have to Multitask in This Economy](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JacquelinePangPaulWoringer.pdf)  | Jacqueline Pang, Paul Woringer  |  
| [BERTogether: Multitask Ensembling with Hyperparameter Optimization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ErikLunaIvanMirandaLiongson.pdf)  | Erik Luna, Ivan Miranda Liongson  |  
| [Implementation of BERT with Projected Attention Layers and Its Effectiveness](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DayoungKimWanbinSong.pdf)  | Dayoung Kim, Wanbin Song  |  
| [Experiments in Improving NLP Multitask Performance](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CarlShan.pdf)  | Carl Shan  |  
| [ExTraBERT: Exclusive Training for BERT Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChinmayKeshavaLalgudiMedhanieIsaiasIrgau.pdf)  | Chinmay Keshava Lalgudi, Medhanie Isaias Irgau  |  
| [Enhancing MinBert Embeddings for Multiple Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DonaldStephens.pdf)  | Donald Stephens  |  
| [minBERT: Contrastive Learning Method](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LongDPham.pdf)  | Long D Pham  |  
| [An examination of multitask training strategies for different BERT downstream tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BjornEngdahlMatthiasHeubi.pdf)  | Bjorn Engdahl, Matthias Heubi  |  
| [Beyond Fine-tuning: Iterative Ensemble Strategies for Enhanced BERT Generalizability](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MeganDassRiyaDulepetShreyaDSouza.pdf)  | Megan Dass, Riya Dulepet, Shreya D'Souza  |  
| [Progressive Layer Sharing on BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NathanielThomasGrudzinski.pdf)  | Nathaniel Thomas Grudzinski  |  
| [Improving minBERT and Its Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MadhumitaVijayDangeYuwenYang.pdf)  | Madhumita Vijay Dange, Yuwen Yang  |  
| [GradAttention: Attention-Based Gradient Surgery for Multitask Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnilYildiz.pdf)  | Anil Yildiz  |  
| [OptiMinBERT: A Comparative Study on the Efficacy of Multitask Versus Specialist Neural Networks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ParasMalhotra.pdf)  | Paras Malhotra  |  
| [Regular(izing) BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EricZhuParkerThomasKasiewicz.pdf)  | Eric Zhu, Parker Thomas Kasiewicz  |  
| [Margin for Error: Exploration of a Dynamic Margin for Cosine-Similarity Embedding Loss and Gradient Surgery to Enhance minBERT on Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexKwonJimmingHe.pdf)  | Alex Kwon, Jimming He  |  
| [MinBERT Task Prioritization, Cross-Attention and Other Extensions for Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArmandoAlejandroBordaParkerJosephStewart.pdf)  | Armando Alejandro Borda, Parker Joseph Stewart  |  
| [Less is More: Exploring BERT and Beyond for Multitask Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LiuxinYangYichunQian.pdf)  | Liuxin Yang, Yichun Qian  |  
| [Exploring LoRA Adaptation of minBERT Model on Downstream NLP Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JamesJosephHennessySuxiLi.pdf)  | James Joseph Hennessy, Suxi Li  |  
| [SMART Multitask MinBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/WeilunChen.pdf)  | Weilun Chen  |  
| [MinBERT and Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/WenlongJi.pdf)  | Wenlong Ji  |  
| [Learning with PALs: Enhancing BERT for Multi-Task Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MichaelQuiSungHoang.pdf)  | Michael Qui Sung Hoang  |  
| [Improving minBERT Embeddings Through Multi-Task Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RahulThapaRohitKhurana.pdf)  | Rahul Thapa, Rohit Khurana  |  
| [BERTina Aguilera: Extensions in a Bottle](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KokhinurKalandarovaMharEisenSantosTenorioSamPrietoSerrano.pdf)  | Kokhinur Kalandarova, Mhar Eisen Santos Tenorio, Sam Prieto Serrano  |  
| [Implementation of minBERT and contrastive learning to improve Sentence Embeddings](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AkshitGoelLinyinLyuNouryaACohen.pdf)  | Akshit Goel, Linyin Lyu, Nourya A Cohen  |  
| [Finetuning minBERT for Downstream Tasks with Multitasking](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NiallThomasKehoePranavSaiRavella.pdf)  | Niall Thomas Kehoe, Pranav Sai Ravella  |  
| [MultiBERT: Enhanced Multi-Task Fine-Tuning on minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChristinaTsangouri.pdf)  | Christina Tsangouri  |  
| [Enhanced Sentence we Embeddings with SimCSE](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BrendanLeeAdamsMcLaughlinChristoDimitrovHristovWilliamShaneHealy.pdf)  | Brendan Lee Adams McLaughlin, Christo Dimitrov Hristov, William Shane Healy  |  
| [A Bilingual BERT Model Ensemble for English-based Multitask Fine-tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RyanJamesDwyer.pdf)  | Ryan James Dwyer  |  
| [Pretrain and Fine-tune BERT for Multiple NLP Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MenggePuYawenGuo.pdf)  | Mengge Pu, Yawen Guo  |  
| [Combining Contrastive Learning with Adaptive Attention and Experimental Dropout to Improve mini-BERT Performance](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JaneneRachanaKimLucyZimmermanRachelLiu.pdf)  | Janene Rachana Kim, Lucy Zimmerman, Rachel Liu  |  
| [Minhbert](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MinhVu.pdf)  | Minh Vu  |  
| [Implementing RO-BERT from Scratch: A BERT Model Fine-Tuned through Regularized Optimization for Improved Performance on Sentence-Level Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CatGonzalesFergesenClarisseYuHokia.pdf)  | Cat Gonzales Fergesen, Clarisse Yu Hokia  |  
| [UmBERTo: Enhancing Performance in NLP Tasks through Model Expansion, SMARTLoss, and Ensemble Techniques](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JulianRodriguezCardenasMayLevin.pdf)  | Julian Rodriguez Cardenas, May Levin  |  
| [Effects of Appropriate Modeling of Tasks and Hyperparameters on Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AdrianLGamarraLafuenteAviUdash.pdf)  | Adrian L Gamarra Lafuente, Avi Udash  |  
| [MiniBERT: Training Jointly on Multiple Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ManasvenGroverXiyuanWang.pdf)  | Manasven Grover, Xiyuan Wang  |  
| [Finetune minBERT for Multi-Tasks Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/YingboLi.pdf)  | Yingbo Li  |  
| [BERTology: Improving Sentence Embeddings for Multi-Task Success](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KyuilLee.pdf)  | Kyuil Lee  |  
| [Not-So-SMART BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EliotKrzysztofJones.pdf)  | Eliot Krzysztof Jones  |  
| [Multi-Tasking BERT: The Swiss Army Knife of NLP](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EstebanWanhoeWuNicoleGarciaSimbaXu.pdf)  | Esteban Wanhoe Wu, Nicole Garcia, Simba Xu  |  
| [Exploring LSTM minBERT with DCT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BenitaWongTinaWu.pdf)  | Benita Wong, Tina Wu  |  
| [BERT Extension Using Sentence-BERT for Sentence Embedding](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnicetDushimeWaMungu.pdf)  | Anicet Dushime Wa Mungu  |  
| [Extending Min-BERT for Multi-Task Prediction Capabilities](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GraceYangXianchenYang.pdf)  | Grace Yang, Xianchen Yang  |  
| [Multitask Learning for BERT Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChunweiChanShuojiaFu.pdf)  | Chunwei Chan, Shuojia Fu  |  
| [Optimizing minBERT for Downstream Tasks using Multitask Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CarlosEmmanuelleAyalaBellido.pdf)  | Carlos Emmanuelle Ayala Bellido  |  
| [Enhancing BERT through Multitask Fine-Tuning, Multiple Negatives Ranking and Cosine-Similarity](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EmilyBroadhurstMichaelMaffezzoli.pdf)  | Emily Broadhurst, Michael Maffezzoli  |  
| [minBERT and Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/QianZhong.pdf)  | Qian Zhong  |  
| [Enhancing Multi-Task Learning on BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ParisZhangYimingNi.pdf)  | Paris Zhang, Yiming Ni  |  
| [BERT’s Odyssey: Enhancing BERT for Multifaceted Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaomingZouMingheZhang.pdf)  | Haoming Zou, Minghe Zhang  |  
| [Efficient Multi-Task MinBERT for Three Default Tasks and Question Answering](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/FanglinLuGerardusdeBruijnRachelRuijiaYang.pdf)  | Fanglin Lu, Gerardus de Bruijn, Rachel Ruijia Yang  |  
| [BERT but BERT-er](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HamzahDaud.pdf)  | Hamzah Daud  |  
| [BERT Mastery: Explore Multitask Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChuLinFuhuXiao.pdf)  | Chu Lin, Fuhu Xiao  |  
| [minBERT using PALs with Gradient Episodic Memory](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChristopherNguyen.pdf)  | Christopher Nguyen  |  
| [Jack of All Trades, Master of Some: Improving BERT for Multitask Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChijiokeMgbahurikeIddahMlauziKwameOcran.pdf)  | Chijioke Mgbahurike, Iddah Mlauzi, Kwame Ocran  |  
| [Multi-task Learning and Fine-tuning with BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MelGuo.pdf)  | Mel Guo  |  
| [Enhanced TreeBERT: High-Performance, Computationally Efficient Multi-Task Model](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/PannSripitakThanawanAtchariyachanvanit.pdf)  | Pann Sripitak, Thanawan Atchariyachanvanit  |  
| [The Best of BERT Worlds: Improving minBERT with multi-task extensions](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JuliusHillebrand.pdf)  | Julius Hillebrand  |  
| [Improving BERT – Lessons from RoBERTa](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChanningLeeHannahGailPrausnitzWeinbaumHaomingSong.pdf)  | Channing Lee, Hannah Gail Prausnitz-Weinbaum, Haoming Song  |  
| [A Rigorous Analysis on Bert’s Language Capabilities](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GauravKiranRane.pdf)  | Gaurav Kiran Rane  |  
| [minBERT and Multitask Learning Enhancements](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NamanGovil.pdf)  | Naman Govil  |  
| [Multi-BERT: A Multi-Task BERT Approach with the Variation of Projected Attention Layer](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaijingZhang.pdf)  | Haijing Zhang  |  
| [QuarBERT: Optimizing BERT with Multitask Learning and Quartet Ensemble](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CarrieGuErickaLiuZixinLi.pdf)  | Carrie Gu, Ericka Liu, Zixin Li  |  
| [BERT-icus, Transform and Ensemble!](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HelenAprilHeMayaWaleriaCzeneszewSidraNadeem.pdf)  | Helen April He, Maya Waleria Czeneszew, Sidra Nadeem  |  
| [Exploring Improvements on BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/SaraHongSophieWu.pdf)  | Sara Hong, Sophie Wu  |  
| [minBERT and Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ShouzhongShi.pdf)  | Shouzhong Shi  |  
| [Efficient Fine-Tuning of BERT with ELECTRA](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AkshayDevGuptaErikRoziVincentJianlinHuang.pdf)  | Akshay Dev Gupta, Erik Rozi, Vincent Jianlin Huang  |  
| [Implementing BERT for multiple downstream tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HamadMMusa.pdf)  | Hamad M Musa  |  
| [Multi Task Fine Tuning of BERT Using Adversarial Regularization and Priority Sampling](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChloeTrujilloMohsenMahvashmohammady.pdf)  | Chloe Trujillo, Mohsen Mahvashmohammady  |  
| [EquiBERT: (An Attempt At) Equivariant Fine-Tuning of Pretrained Large Language Models](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/PatrickJamesSicurello.pdf)  | Patrick James Sicurello  |  
| ["That was smooth": Exploration of S-BERT with Multiple Negatives Ranking Loss and Smoothness-Inducing Regularization](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JohnnyChangKanuGroverKaushalAtulAlate.pdf)  | Johnny Chang, Kanu Grover, Kaushal Atul Alate  |  
| [Enhancing BERT for Advanced Language Understanding: A Multitask Learning Approach with Task-Specific Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnushaAditiKuppahallyMalaviRavindranZiyueJuliaWang.pdf)  | Anusha Aditi Kuppahally, Malavi Ravindran, Ziyue (Julia) Wang  |  
| [Triple-Batch vs Proportional Sampling: Investigating Multitask Learning Architectures on minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EthanSargoTiaoRikhilPareshVagadia.pdf)  | Ethan Sargo Tiao, Rikhil Paresh Vagadia  |  
| [Extending Applications of Layer Selecting Rank Reduction](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AbrahamAlappat.pdf)  | Abraham Alappat  |  
| [Improving BERT for Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KevinNguyenPhan.pdf)  | Kevin Nguyen Phan  |  
| [BEAKER: Exploring Enhancements of BERT through Learning Rate Schedules, Contrastive Learning, and CosineEmbeddingLoss](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ElizabethTheresaBaena.pdf)  | Elizabeth Theresa Baena  |  
| [Gradient Descent in Multi-Task Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/DavidSaykinKfirShmuelDolev.pdf)  | David Saykin, Kfir Shmuel Dolev  |  
| [SMART Surgery: Combining Finetuning Methods for Multitask BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/EthanPaulFoster.pdf)  | Ethan Paul Foster  |  
| [Fine-tuning minBERT For Multi-task Classification](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JimmyOtienoOgada.pdf)  | Jimmy Otieno Ogada  |  
| [Three Headed Mastery: minBERT as a Jack of All Trades in Multi-Task NLP](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IfditaHasanOrneyRafaelPerezMartinezValerieAnnFanelle.pdf)  | Ifdita Hasan Orney, Rafael Perez Martinez, Valerie Ann Fanelle  |  
| [Balancing Performance and Computational Efficiency: Exploring Low-Rank Adaptation for Multi-Transferring Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/CarolineSantosMarquesdaSilva.pdf)  | Caroline Santos Marques da Silva  |  
| [ExtraBERT: Applying BERT to Multiple Downstream Language Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IsaacIGorelikRishiDange.pdf)  | Isaac I. Gorelik, Rishi Dange  |  
| [BERT and Beyond: A Study of Multitask Learning Strategies for NLP](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/FebieJaneLinJackPLe.pdf)  | Febie Jane Lin, Jack P Le  |  
| [Mini Bert Optimized for Multi Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LinLin.pdf)  | Lin Lin  |  
| [Methods to Improve Downstream Generalization of minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/RamgopalVenkateswaran.pdf)  | Ramgopal Venkateswaran  |  
| [Maximizing MinBert for Multi-Task Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JordanAndyParedesShumannRXu.pdf)  | Jordan Andy Paredes, Shumann R Xu  |  
| [minBERT Multi-Task Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AntonioDaviMacedoCoelhodeCastro.pdf)  | Antonio Davi Macedo Coelho de Castro  |  
| [Fine-tuning minBERT for multi-task prediction](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/IshitaMangla.pdf)  | Ishita Mangla  |  
| [Grid Search for Improvements to BERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HarshGoyal.pdf)  | Harsh Goyal  |  
| [SMART loss vs DeBERTa](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/MichaelLiuMichaelPhillipHayashiRobertoLobatoLopez.pdf)  | Michael Liu, Michael Phillip Hayashi, Roberto Lobato Lopez  |  
| [Extending Phrasal Paraphrase Classification Techniques to Non-Semantic NLP Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/NikhilSharmaSamyCherfaoui.pdf)  | Nikhil Sharma, Samy Cherfaoui  |  
| [Loss Weighting in Multi-Task Language Learning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JuliaKwak.pdf)  | Julia Kwak  |  
| [Task-specific attention](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChaoqunJia.pdf)  | Chaoqun Jia  |  
| [minBERT and Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/XinpeiYu.pdf)  | Xinpei Yu  |  
| [Integrating Cosine Similarity into minBERT for Paraphrase and Semantic Analysis](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GeraldJohnSufleta.pdf)  | Gerald John Sufleta  |  
| [SlapBERT: Shared Layers and Projected Attention For Enhancing Multitask Learning with minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AlexHeZhaiAllisonJiaDevenKiritPandya.pdf)  | Alex He Zhai, Allison Jia, Deven Kirit Pandya  |  
| [MT-DNN with SMART Regularisation and Task-Specific Head to Capture the Pairwise and Contextually Significant Words Interplay](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/HaoyuWang.pdf)  | Haoyu Wang  |  
| [Speedy SBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/LeythRamezToubassyReneeDuarteWhite.pdf)  | Leyth Ramez Toubassy, Renee Duarte White  |  
| [Using Gradient Surgery, Cosine Similarity, and Additional Data to Improve BERT on Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ChanseHBhaktaJosephAnthonySeibaKasenStephensen.pdf)  | Chanse H. Bhakta, Joseph Anthony Seiba, Kasen Stephensen  |  
| [Multitask BERT Model with Regularized Optimization and Gradient Surgery](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/JennyXu.pdf)  | Jenny Xu  |  
| [BitBiggerBERT: An Extended BERT Model with Custom Attention Mechanisms, Enhanced Fine-Tuning, and Dynamic Weights](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/KhanhVTranThomasCharlesHatcherVladimirAGonzalezMigal.pdf)  | Khanh V Tran, Thomas Charles Hatcher, Vladimir A Gonzalez Migal  |  
| [minBERT and Downstream Tasks Final Report](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/BingqingZuYixuanLin.pdf)  | Bingqing Zu, Yixuan Lin  |  
| [Evaluating Contrastive Learning Strategies for Enhanced Performance in Downstream Tasks](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GeorgiosChristoglouZacharyEvansBehrman.pdf)  | Georgios Christoglou, Zachary Evans Behrman  |  
| [A SMARTer minBERT](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/ArisaSugiyamaChueDaphneLiuPoonamSahoo.pdf)  | Arisa Sugiyama Chue, Daphne Liu, Poonam Sahoo  |  
| [OptimusBERT: Exploring BERT Transformer with Multi-Task Fine-Tuning, Gradient Surgery, and Adaptive Multiple Negative Rank Loss Learning Fine-Tuning](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/GabeEduardoSeirRyderThompsonMathenyShawnCharles.pdf)  | Gabe Eduardo Seir, Ryder Thompson Matheny, Shawn Charles  |  
| [Optimizing minBert via Cosine Similarity and Negative Sampling](https://web.stanford.edu/class/archive/cs/cs224n/cs224n.1244/final-projects/AnanyaSiriVasireddyNehaVinjapuri.pdf)  | Ananya Siri Vasireddy, Neha Vinjapuri  |
