v0.39.1 Now with live service preview

Get the job done without burning through your wallet.

From plan to pull request, Ogcode does the work end to end. OGmap steers it straight to the right files in even a huge codebase, it reads only the lines that matter, and it recalls instead of re-sending your whole conversation every turn — so you pay about a third of the usual bill.

Star on GitHub

Recall, not replay.

Every decision, every dead end, every why — indexed once, recalled on demand. Only what this turn needs is sent.

Built for developers

Ship with an agent that already knows the repo. It plans features, opens PRs, and never asks you to re-explain the codebase.

Built for long projects

Month three feels like day one. Conventions, choices and rejected ideas carry over — no amnesia.

Built for your machine

One Go binary, self-hosted, AGPL. No account, no cloud. Pair it with Ollama and no model call leaves the room.

A small context. Sharp models.

Ogcode sends only what the current turn needs, so requests stay well under 200K tokens however long a session runs. Models read a lean context most accurately, which is why smaller, cheaper and local models do their best work here.

One real session GLM-5.3 Flash · 42 hours · 2,589 steps
Tokens sent with each request, cached input included. Half carried under 54K, the busiest never passed 168K, and the context stayed level from the first hour to the last. The model’s window is 1M; the session used a sixth of it at most.

54K

Typical request

The median across 21,252 requests in 449 sessions.

99%

Under 180K

Almost every request leaves the model room to think.

168K

Longest session’s peak

2,589 steps, and the context never crept up.

Accuracy holds up

Models read a short, focused prompt more reliably than a long one. Details buried deep in a huge context are the first to be missed, so a lean request keeps every answer grounded.

Small models keep up

Flash, mini and local models that stumble on a bloated context do solid work on a lean one. Pick the cheaper model and still get the job done.

Your model. Your keys.

Or none at all. Bring a key from any major provider — or run fully local with Ollama. Switching models never resets what Ogcode remembers.

Claude gets a tight, recalled context every turn — not a transcript that grows until it breaks.

setup export ANTHROPIC_API_KEY=sk-ant-… && ogcode Docs on configuring model providers

One OpenAI-compatible slot points at GPT — or at Gemini, Groq, DeepSeek or any provider on the belt below.

setup export OPENAI_API_KEY=sk-… && ogcode Docs on configuring model providers

Hundreds of models behind a single key. Change models mid-project; the memory stays with the project.

setup export OPENROUTER_API_KEY=sk-or-… && ogcode Docs on configuring model providers

Local and free. Ogcode auto-detects a running Ollama, so your prompts stay on your machine.

setup ollama serve & ogcode Docs on configuring model providers

Rather pay one flat price? OGX: $2 of tokens every day, for $5 a month

And any OpenAI-compatible endpoint…

Or use ours. Hosted in-house.

OGX runs DeepSeek and GLM on our own high-performance inference cluster. Replies start in about a second and stream faster than you can read them.

OGX · in-house inference cluster DeepSeek V4.1 Flash · GLM 5.3 Flash
Both models run on hardware we operate ourselves, so the speed you see and the uptime you get are ours to answer for.

about 310tok/s

DeepSeek V4.1 Flash

Median output speed on real coding prompts.

about 215tok/s

GLM 5.3 Flash

Median output speed on the same prompts.

about 1s

To first token

Replies begin about a second after you send, on both models.

$5/ month

One flat price

Both models come with OGX: $2 of tokens every day, $60 a month, and a bill that never moves.

See the OGX plan

Why Ogcode

Everything a long session needs to stay sharp — and cheap.

Replaying the whole chat

Without recall, every turn re-sends the entire conversation. The context window fills up, the bill climbs with every message, the oldest turns fall out of the window, and the same files are read in full again and again.

Recalling with Ogcode

With Ogcode, each turn recalls only what it needs — the relevant decisions and the exact file ranges — so every turn sends one small, tight context window, and what you decided months ago is still remembered.

one tight window, every turn

Pay only for what the turn needs

(Not for the last two hundred turns, again.)

The longer you work, the more you save.

Other agents re-send the whole conversation every turn, so the bill climbs with every message. Ogcode sends one tight window, so it stays flat.

See the numbers

No context clock.

Long sessions stop being a countdown to amnesia. The window carries this turn's context — not every turn that came before it.

How it works

Day 90, it still remembers.

Decisions, conventions and rejected approaches are written down as the work happens — and brought back the moment they matter again.

What it remembers

Same model, sharper answers.

Other agents hand the model every old turn, stale tool output and whole file, then ask it to find the answer in the noise. Ogcode's context is garbage-free — only the decisions, files and line ranges this question needs — so answers land on point, and the first attempt is usually the review-ready one.

How it works

Ogcode

the agent itself

$0 /forever

Read the license

Hosted models

bring your own key

−70% tokens

vs. full-transcript replay

See the numbers

Local models

with Ollama

$0 /offline

Run it locally

Your server

for teams

Self-host

behind your own auth

Remote deployment

Context engineering, at its peak.

Ogcode maps before it reads, reads ranges instead of whole files, and recalls instead of replaying. Every query gets its own context, built on the fly.

Plan Mode ships pull requests

Describe the outcome and lock the plan. Ogcode breaks it into a board of tasks, runs independent ones in parallel in isolated git worktrees, and opens a pull request for each.

where do retries live? retry.go api.go policy.ts

Powered by OGmap

Think Google Maps, but for your agent. OGmap charts your whole codebase — every module, route and helper — so the agent always knows where things live and heads straight there.

No more blind file reads

Every declaration is mapped to its lines, so the agent reads the exact range a task needs — and leaves the other four hundred lines alone.

Context built on the fly

Ogcode doesn't drag a fixed history into every turn. For any query, it builds the context it needs on the spot — recalling past decisions, finding the right files with OGmap, and reading just the lines that matter. With less noise to fight through, the first attempt is usually the review-ready one.

Your project is more than code

Specs, proposals and spreadsheets sit right beside the source. Ogcode reads PDF, DOCX and CSV natively — no exports, no copy-paste.

A race car at full sprint with light trails, beside a glowing speedometer.

Fast, even on a monorepo

Reads only the lines a task needs, so turns finish in seconds, not minutes — and stay that fast on a ten-year-old monorepo.

Nothing goes unreviewed.

Other agents commit in the dark. Ogcode works in the open — in Build Mode every change waits in your working tree as a real git diff, and nothing is written to history until you say so.

A real diff, not a done deal

Every change lands in your working tree. Inspect it, stage it, unstage it, or throw it away.

Hunk by hunk

Review file by file and hunk by hunk before anything touches history.

A human says yes

Nothing merges without a human yes.

Fair questions.

The ones skeptics ask first — answered plainly.

Does my source code leave my machine?
Only in the prompts you send to whichever model provider you chose — the same as any coding agent. There is no Ogcode server, no telemetry endpoint, and no account. Everything lives in a local database. Pick a local Ollama model and nothing leaves the machine at all.
Can I really run it fully offline?
Yes. It auto-detects a running ollama on macOS and Linux. Offline you lose the web search tool and hosted models — everything else works.
Is the 70% number real, or marketing?
Measured against agents that re-send the whole conversation every turn, across real sessions — about 68% at 50 messages, ~76% at 1,000. Short sessions save less because there's less history to avoid re-sending. The mechanism and the numbers are both in the README, and the code is open source — check it yourself.
If it doesn't re-read everything, does it forget things?
It doesn't drop them — it moves them. What matters is remembered and brought back when it's relevant. In practice it forgets less than other agents, which eventually run out of window and start losing the beginning of the conversation.
Can a small, cheap model really keep up?
Far better than in an agent that lets the context balloon. Ogcode keeps each request lean: across 449 sessions the median request carried 54K tokens and 99% stayed under 180K. Models, small ones most of all, answer most accurately when the prompt is short and focused. Our own longest session ran 2,589 steps on GLM-5.3 Flash and never carried more than 168K. See the chart
Does it work on Windows?
Yes — irm https://ogcode.xyz/install.ps1 | iex in PowerShell, or run the Docker image. Opening PRs automatically needs the GitHub CLI authenticated.
Can I run it on a server and use it from my laptop?
Yes. It's a normal web server — put it behind nginx or Caddy with a password. Don't expose it to the open internet unprotected, since it can run commands on your machine. Setup notes are in the remote deployment guide.
What does it cost?
Ogcode is free and open source under the AGPL-3.0. You pay your model provider for tokens — which is rather the point of sending 70% fewer of them. With Ollama, it costs nothing.
I already use Claude Code / Cursor. Why switch?
You may not need to. Ogcode is strongest when sessions run long, when you want a feature planned and shipped as PRs rather than a file edited, or when you want to stay off a subscription and off the cloud. If you mostly want inline completions in your editor, keep what you have.
How is this different from just a bigger context window?
A bigger window still re-sends everything, every turn — you pay for it either way, and models get less accurate as the prompt grows. Ogcode brings back only what's relevant, so requests stay well under 200K and the bill stays flat no matter how long you work.

Still have a question?

Ask in Discord or open an issue — a human answers.

  • AGPL-3.0
  • one binary

From side project to enterprise codebase.

One Go binary. Any model, your keys. No account and no cloud — your sessions and memory stay in a local database.

Read the docs
  • One Go binary
  • No account, no cloud
  • Any model, your keys
  • Runs fully offline with Ollama
  • Recalls instead of re-sending
  • Reads PDF, DOCX and CSV
  • Every change is a reviewable git diff
  • Plan Mode opens pull requests
  • Open source, AGPL-3.0