Based on the available codebase memory context, I cannot determine how asyncio semaphore specifically controls Ollama concurrency in local mode. The retrieved entities show: 1. `_should_use_cloud` - Determines cloud usage #main.py:986 2. `MEMORY_MODE` configuration with cloud/hybrid modes #config.py:30-45 3. `_background_scan` - Uses Semaphore(4) for tree-sitter and Semaphore(2) for LLM summarization during scanning #main.py:1827 None of these entities contain the specific concurrency logic for `_query_ollama` at runtime. If this routing exists in the codebase, it's not present in the retrieved memory entries.