How to set up a local & offline coding assistant
This guide sets up a local coding assistant inside your editor. If you instead want to run a local model and call it from your own app, see Build a Local LLM App, and Cloud vs Local Models for when local makes sense.
Local models are viable now
As of mid-2026, local models have closed much of the gap with frontier APIs for day-to-day coding tasks. Vicki Boykis documents this well in Running local models is good now (2026-06-15): on an M2 Mac with 64 GB RAM she runs agentic coding workflows -- refactoring notebooks, generating unit tests, bootstrapping repos -- at roughly ~75% the accuracy and speed of frontier models, without a cloud API call. For agentic coding she highlights Gemma 4 26B A4B and Gemma 4 12B QAT; she has also used Qwen 3 MOE and Qwen 2.5 Coder, which remain practical local options for this setup.
The remaining limitations are real: inference is slower than a remote API, context windows are capped by your RAM, and results still warrant a second opinion on tricky problems. But for privacy-sensitive work or offline development, the tooling is now good enough to use daily.
Ollama + Continue Dev
- Install ollama
- Install the continue.dev Plugin for JetBrains or VSCode
- Optionally, use OpenWebUI via Docker as an Interface for Chatting
Keep Ollama, LM Studio, and Open WebUI bound to localhost unless you add authentication and network controls.
Ollama binds to 127.0.0.1:11434 by default, but OLLAMA_HOST=0.0.0.0:11434 exposes it. LM Studio API
tokens are off by default. Open WebUI authentication is on by default through WEBUI_AUTH=True; do not set
WEBUI_AUTH=False on a shared or exposed instance. See the full security note in
Build a Local LLM App.
Coding assistants can fill context quickly with files, diffs, and repo maps. Ollama's effective context window
now depends on version and VRAM; configure it with OLLAMA_CONTEXT_LENGTH, API num_ctx, or a Modelfile only
after checking memory. See Build a Local LLM App and
Serving LLMs at Scale.
Model Setup
- Quick Tab Completion
ollama pull qwen2.5-coder:1.5b
- Indexing and Codebase Search
ollama pull nomic-embed-text
- General Purpose Reasoning Model
ollama pull phi4
- https://ollama.com/library/phi4
- MIT License
- Alternatively use Gemma 4 (strong mid-2026 recommendation, especially on Apple Silicon with 32 GB+ RAM):
- https://ollama.com/library/gemma4
gemma4:26bfor the 26B variant,gemma4:12bfor the lighter 12B variant
ollama pull gemma4:26b# or the lighter variantollama pull gemma4:12b - Or Qwen 3 MOE -- mixture-of-experts, punches above its weight on coding:
ollama pull qwen3:30b-a3b
- Optional reranking model
- Continue's reranking docs document Voyage, Cohere, LLM-based reranking, and Hugging Face Text Embeddings Inference (TEI).
- They explicitly warn that the LLM fallback does not work with local models such as Ollama because too many parallel requests are required.
- For a local setup, run a TEI reranker separately and use the
huggingface-teiconfig shown below, or omit the reranker until you have a measured need for it.
- Update continue.dev
config.yaml-> see here - Run ollama api locally
ollama serve
Open Web UI

The following Docker commands and Compose example publish Open WebUI to 127.0.0.1. On Docker Engine
versions older than 28.0.0, hosts on the same layer-2 network segment may still reach ports published to
localhost, so this binding does not guarantee host-only access. Use a firewall or other network controls
on those versions. See Docker's port-publishing documentation.
Nvidia GPU
docker run -d -p 127.0.0.1:3000:8080 --gpus all --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart unless-stopped ghcr.io/open-webui/open-webui:cuda
Other
docker run -d -p 127.0.0.1:3456:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart unless-stopped ghcr.io/open-webui/open-webui:main
Docker Compose
Below docker connects to ollama running natively on windows and not via docker.
services:
open-webui:
image: ghcr.io/open-webui/open-webui:cuda
container_name: open-webui
volumes:
- ./data:/app/backend/data
ports:
- 127.0.0.1:3456:8080
environment:
- 'OLLAMA_BASE_URL=http://host.docker.internal:11434'
extra_hosts:
- host.docker.internal:host-gateway
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [ gpu ]
volumes:
open-webui: { }
The main and cuda tags move over time. They are convenient for a local experiment, but pin a tested
release tag or image digest before using this setup for work you need to reproduce. Back up the persistent
data volume before upgrades.
Usage
Use directly in your editor

or via the chat-sidebar tab

Suggested continue.dev config
- Unix:
~/.continue/config.yaml - Windows:
%USERPROFILE%\.continue\config.yaml
Continue's current codebase/documentation awareness guide says the old @Codebase and @Docs context
providers are deprecated in favor of Agent mode tools, project rules, and MCP servers. Continue's deprecated
codebase reference also covers @Folder. Use built-in file/search/repo-map tools and .continue/rules for
project context; use MCP servers such as Context7 or custom internal docs servers when documentation retrieval
must be tool-backed.
name: Local Ollama
version: 0.0.1
schema: v1
models:
- name: Gemma 4 26B
provider: ollama
model: gemma4:26b
roles:
- chat
- edit
- apply
- name: Qwen 3 MOE
provider: ollama
model: qwen3:30b-a3b
roles:
- chat
- edit
- apply
- name: Phi-4
provider: ollama
model: phi4
roles:
- chat
- edit
- apply
- name: Qwen2.5-Coder
provider: ollama
model: qwen2.5-coder:1.5b
roles:
- autocomplete
- name: Nomic Embed Text
provider: ollama
model: nomic-embed-text
roles:
- embed
- name: BGE Reranker via TEI
provider: huggingface-tei
model: tei
apiBase: http://localhost:8080
apiKey: tei
roles:
- rerank
rules:
- Always respond in English.
- Always provide clear, concise, and accurate answers.
- Support a software developer with explanations, code, best practices, and debugging help.
prompts:
- name: test
description: Generate unit tests for the highlighted code.
prompt: |
Write a comprehensive set of unit tests for the provided code. Ensure to include setup, execution of correctness checks with important edge cases, and teardown. Present the tests as plain text output.
- name: refactor
description: Improve the code's structure for better readability.
prompt: |
Refactor the provided code to improve its structure and readability without altering its functionality. Include a detailed explanation of your changes and reasoning.
- name: optimize
description: Enhance code performance with a detailed explanation of changes and trade-offs.
prompt: |
Optimize the provided code for performance while maintaining its current behavior. Describe any trade-offs involved in your optimization process.
- name: explain
description: Analyze and explain the code's functionality and potential improvements.
prompt: |
Explain the logic and functionality of the provided code. Discuss any potential inefficiencies or unnecessary computations that could be improved for better performance.
- name: document
description: Create clear and concise function documentation using the correct language format.
prompt: |
Write language-specific documentation for the provided function. Use appropriate formats like Javadoc for Java or JSDoc for JavaScript. Ensure clarity and conciseness in your explanation.
context:
- provider: diff
- provider: file
- provider: code
- provider: currentFile
- provider: terminal
- provider: open
- provider: repo-map
- provider: tree
- provider: problems
- provider: os
- provider: web
- provider: url
- provider: docs # legacy @Docs; see the deprecation note above
docs:
- name: aem.live
startUrl: https://www.aem.live/docs
favicon: https://www.aem.live/favicon.ico
- name: AEMaaCS
startUrl: https://experienceleague.adobe.com/de/docs/experience-manager-cloud-service
favicon: https://experienceleague.adobe.com/favicon.ico
- name: lucanerlich
startUrl: https://lucanerlich.com
- name: react
startUrl: https://react.dev/
- name: typescript
startUrl: https://www.typescriptlang.org/
- name: react spectrum
startUrl: https://react-spectrum.adobe.com/index.html
See also
- Build a Local LLM App - Ollama, LM Studio, Open WebUI, and local API security
- Cloud vs Local Models - choosing local, self-hosted, or managed models
- Serving LLMs at Scale - context length, KV cache, batching, and serving trade-offs
- Multimodal & Voice - images, documents, audio, and realtime voice agents