| | server { | | | listen 80; | | | server_name ip address; | | | access_log /var/log/nginx/access.log; | | | error_log /var/log/nginx/error.log; | | | root /var/www/src/api/static; | | | | | | location / { | | | include uwsgi_params; | | | uwsgi_pass unix:/var/www/viberary.sock; | | | proxy_pass http://127.0.0.1:8000; | | | } | | | } | | | ` | | | | | | # Install docker | | | sudo apt install apt-transport-https ca-certificates curl software-properties-common | | | curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo apt-key add - | | | sudo add-apt-repository "deb [arch=amd64] https://download.docker.com/linux/ubuntu focal stable" | | | apt-cache policy docker-ce | | | sudo apt install docker-ce | | | sudo systemctl status docker | | | sudo apt-get install docker-compose-plugin | | | | | | | | | # set up app | | | make build | | | | | | #ssh over transformer model | | | scp | | | make up-intel | | | make embed | | | | | | voila | | | | | | # Metrics and alerting agent | | | curl -sSL https://repos.insights.digitalocean.com/install.sh | sudo bash | | | ps aux | grep do-agent | [view raw](https://gist.github.com/veekaybee/f5ff921355e6cd3970bd097dcb0fbc35/raw/105e3210fb3154323451d3d000a69008d48ff912/spin_up.sh) [ spin_up.sh ](https://gist.github.com/veekaybee/f5ff921355e6cd3970bd097dcb0fbc35#file-spin_up-sh) hosted with ❤ by [GitHub](https://github.com) Finally everything is routed to port 80 via nginx, which I configured on each DigitalOcean droplet that I created. I load balanced two droplets behind a load balancer, pointing to the same web address, a domain I bought from Amazon’s Route 53. I eventually had to transfer the domain to Digital Ocean, because it’s easier to manage SSL and HTTPS on the load balancer when all the machines are on the same provider. This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars) [ Show hidden characters ](https://vickiboykis.com/2024/01/05/retro-on-viberary/{{%20revealButtonHref%20}}) | | server { | | --- | --- | | | listen 80; | | | server_name ip address; | | | access_log /var/log/nginx/access.log; | | | error_log /var/log/nginx/error.log; | | | root /var/www/src/api/static; | | | | | | location / { | | | include uwsgi_params; | | | uwsgi_pass unix:/var/www/viberary.sock; | | | proxy_pass http://127.0.0.1:8000; | | | } | | | } | [view raw](https://gist.github.com/veekaybee/f18ce09aa50c7cfdcb61300770ef8f52/raw/d0a2a08d401d1bb1196c7c14eba95e49d0d5a9f5/nginx) [ nginx ](https://gist.github.com/veekaybee/f18ce09aa50c7cfdcb61300770ef8f52#file-nginx) hosted with ❤ by [GitHub](https://github.com) Now, we have a working app. The final part of this was load testing, which I did with [Python’s Locust library](https://locust.io/), which provides a nice interface for running any type of code against any endpoint that you specify. One thing that I realized as I was load testing was that my model was slow, and search expects instant results, so I converted it to an [ONNX artifact](https://blog.vespa.ai/stateful-model-serving-how-we-accelerate-inference-using-onnx-runtime/) and had to change the related code, as well. ![](https://vickiboykis.com/images/locust.jpeg) Finally, I wrote a small logging module that propogates across the app and keeps track of everything in the docker compose logs. # Key Takeaways * **Getting to a testable prototype is key**. I did all my initial exploratory work locally in Jupyter notebooks, [including working with Redis](https://github.com/veekaybee/viberary/blob/main/src/notebooks/05_duckdb_0.7.1.ipynb), so I could see the data output of each cell. I [strongly believe](https://vickiboykis.com/2021/11/07/the-programmers-brain-in-the-lands-of-exploration-and-production/) working with a REPL will get you the fastest results immediately. Then, when I had a strong enough grasp of all my datatypes and data flow, I immediately moved the code into object-oriented, testable modules. Once you know you need structure, you need it immediately because it will allow you to develop more quickly with reusable, modular components. * **Vector sizes and models are important**. If you don’t watch your hyperparameters, if you pick the wrong model for your given machine learning task, the results are going to be bad and it won’t work at all. * **Don’t use the cloud if you don’t have to**. I’m using DigitalOcean, which is really, really, really nice for medium-sized companies and projects and is often overlooked over AWS and GCP. I’m very versant in cloud, but it’s nice to not have to use BigCloud if you don’t have to and to be able to do a lot more with your server directly. DigitalOcean has reasonable pricing, reasonable servers, and a few extra features like monitoring, load balancing, and block storage that are nice coming from BigCloud land, but don’t overwhlem you with choices. They also recently acquired [Paperspace](https://www.paperspace.com/), which I’ve used before to train models, so should have GPU integration. * **DuckDB** is becoming a stable tool for work up to 100GB locally. There are a lot of issues that still need to be worked out because it’s a growing project. For example, for two months I couldn’t use it for my JSON parsing because it didn’t have regex features that I was looking for, which were added in 0.7.1, so use with caution. Also, since it’s embedded, you can only run one process at a time which means you can’t run both command line queries and notebooks. But it’s a really neat tool for quickly munging data. * **Docker still takes time** I spent a great amount of time on Docker. Why is Docker different than my local environment? How do I get the image to build quickly and why is my image now 3 GB? What do people do with CUDA libraries (exclude them if you don’t think you need them initially, it turns out). I spent a lot of time making sure this process worked well enough for me to not get frustrated rebuilding hundreds of times. Relatedly, **Do not switch laptop architectures in the middle of a project** . * **Deploying to production is magic** , even when you’re a very lonely team of one, and as such [is filled with a lot of unknown variables](https://vickiboykis.com/2021/06/20/the-ritual-of-the-deploy/), so make your environments as absolutely reproducible as possible. And finally, * True semantic search is very hard and involves a lot of algorithmic fine-tuning, both in the machine learning, and in the UI, and in deployment processes. People have been fine-tuning Google for years and years. Netflix had thousands of labelers. [Each company has teams of engineers working on search and recommendations](https://vicki.substack.com/p/what-we-talk-about-when-we-talk-about) to steer the algorithms in the right direction. Just take a look at the company formerly known as Twitter’s algo stack. It’s fine if the initial results are not that great. The important thing is to keep benchmarking the current model against previous models and to keep iterating and keep on building. # Citations