Cool! You can now create your own applications with your own local model quite easily. ··· ## Fine-Tuning the Experience — Taming the “Thinking” Tokens Qwen 3 is a hybrid reasoning model. By default, it generates a verbose `...` block outlining its chain of thought before providing the actual answer. Sometimes you want to see the math but most of the time, you just want the answer quickly (and cut some time from waiting the output tokens from the thinking process). Here is how you bypass the reasoning pass: * **Disable it entirely:** `ollama run qwen3:8b --think=false` * **Run it, but hide it from the UI:** `ollama run qwen3:8b --hidethinking` * **In scripts:** Pass `"think": false` in your JSON payload. ··· ## A Warning About Web Search Models are static up until their training data. That means that they can’t access data after they were trained, and companies have been relying on web search tools to augment the capability of the models. For example for our local model: ![](https://assets.insightmediagroup.io/media/v2/resize:fit:700/1*Domg-UrTms4V4EfpWMevZQ.png) Last day of training data of our Local Model But, Ollama allows you to hand the model a web-search tool. This sounds incredible but there’s a catch. The search itself executes on Ollama’s hosted cloud service. The moment you enable it, your prompts are being sent over the internet to fetch search results. The model stays local, but your queries do not. This may violate the principle of privacy you want to guarantee with the setup. ··· ## Bonus: VS Code Integration The ultimate endgame for me was getting an offline coding assistant. The cleanest, entirely free path for this is the **Continue.dev** extension. * Install VS Code and the Continue extension. * Open Continue’s configuration file at `~/.continue/config.yaml`. * Point it at your local Ollama server: yaml