AI Coding Field Notes

Coding agent pricing comparison: cheapest LLM API for coding, cost per million tokens, from the OpenRouter price list as JSON and CSV. An AI model token price comparison repriced at a real agent mix — 95.6% cached input pricing, 6.5x off list: prompt caching price, cache read price, cache hit share, long context pricing, DeepSeek peak/off-peak.

View the Project on GitHub xyzs996/llm-api-pricing

How Alibaba’s Open Code Review Slashed AI Code Review Costs by 90% — And What It Means for Independent Developers

How Alibaba’s Open Code Review Slashed AI Code Review Costs by 90% — And What It Means for Independent Developers

Written with AI assistance. Figures without a traceable source were cut before publishing.

The same piece is posted as a thread on GitHub — that copy has a reply box under it, and this one does not.

In July 2026, Alibaba open-sourced Open Code Review. Within weeks, independent developers across China tested it on 200 pull requests spanning 50 repositories. The result? Higher accuracy than Claude Code — and a 90% drop in token usage. Where Claude needed ~90K input tokens per PR, Open Code Review used just ~10K. This wasn’t benchmark theater. These were real projects in private GitHub orgs, covering JavaScript, Python, Go, and seven other languages. The tool didn’t just optimize prompts. It rethought the entire architecture of AI-assisted code review.

And so far, this shift has gone largely unnoticed in the West.

What Happened in the Chinese AI Scene That Western Devs Missed

A quiet but consequential shift is underway among Chinese independent developers: they’re moving away from feeding raw diffs into general-purpose LLM agents like Claude and toward hybrid systems that combine deterministic logic with targeted AI calls. Open Code Review exemplifies this. It doesn’t dump your entire PR into the model. Instead, it uses deterministic modules to filter files, parse diffs, match rules, and bundle only the relevant context before invoking an LLM. In one test case, a PR with a large diff would have required over 50K tokens if full context were passed. Open Code Review reduced that to under 10K by excluding unchanged files and non-relevant modules.

The tool supports line-level structured output — think file.js:line: warning: avoid hardcoded API keys — severity levels (info/warning/error), and CI/CD integration via GitHub Actions using a pre-built Docker image tagged ocr-v0.3.1. It also handles session recovery through a local .ocr_cache/ directory, so if your network drops mid-review, you don’t restart from zero. That stability matters. Cloud-only agents often fail silently when latency spikes, leaving PRs half-reviewed.

This isn’t just incremental optimization. It’s a fundamental architectural rethink — one that’s absent from most Western “prompt engineering” discussions — which was directly inspired by Ali Sadeghi’s 7-command pipeline, where each step writes intermediate results to disk in order to prevent the AI from drifting off-spec, a design choice that enforces strict modularity and ensures every LLM call begins with a clean state, thereby eliminating context pollution and aligning the system with deterministic software engineering practices rather than conversational paradigms. In Open Code Review, rules and specs live in files like CODEREVIEW_RULES.md, not buried in system prompts. Every LLM call starts fresh, avoiding context pollution from prior rounds. The result is a system that treats the AI as a component in a larger engineering pipeline, not a copilot you chat with. And that distinction is costing Western devs — literally.

This approach extends beyond code review. Developers are increasingly adopting project-level configuration files such as CLAUDE.md or AGENTS.md to codify collaboration rules and workflows directly within the codebase. By doing so, they turn tacit knowledge into executable skills that any team member—or AI agent—can follow consistently, reducing onboarding time and execution drift. This practice aligns with the broader trend of “externalizing memory” to maintain fidelity across long development cycles.

Another emerging pattern is the use of tools like agent-device, which enables AI agents to interact with mobile apps through CLI commands—opening apps, tapping UI elements, or entering text—by referencing on-screen components. This capability allows developers to automate exploratory testing and smoke checks without writing traditional test scripts, especially when integrated with Codex or Claude Code. For instance, agent-device tap @e2 triggers an action based on a dynamically assigned element ID.

Open Generative AI has gained traction by bundling over 200 models into a self-hosted video generation studio, addressing content moderation, subscription costs, and data control. With 15K GitHub stars and 2.6K forks, it reflects strong developer demand for open, auditable AI tooling that doesn’t lock users into opaque platforms.

Why the Hybrid Architecture Beats Pure LLM Agents

Open Code Review’s core innovation is its staged pipeline. Phase 1 is fully deterministic: file type filtering → change detection → rule matching → context bundling. Only then does Phase 2 invoke the LLM, which receives structured prompts to audit changes against explicit, external specifications. This hybrid architecture of deterministic engineering and targeted LLM use improves accuracy in security anti-pattern detection compared to general-purpose agents like Claude Code. By anchoring the review in fixed rules and pre-processed context, the system ensures consistent, accurate feedback while minimizing unnecessary model calls.

This approach directly addresses two persistent problems in long agent sessions: drift and collusion. “Drift” happens when an AI’s internal state diverges from reality over time. By reading from disk instead of chat history, Open Code Review eliminates this. “Collusion” occurs when an AI generates code and then justifies it using its own flawed logic — a form of self-reinforcing hallucination. One developer reported a case where Claude rewrote a function to match its own outdated explanation. Open Code Review caught it because the spec file hadn’t changed. The “review” role enforces consistency with external truth, not internal coherence.

The token savings are structural, not marginal. Pure LLM agents often re-read project context — package.json, tsconfig.json, utility files — on every turn. As projects age, this cost compounds. One team working on a mature TypeScript codebase observed that their AI assistant consumed a number of tokens across multiple PRs, with the majority of the cost stemming from repeatedly parsing stable configuration files and shared modules. In contrast, Open Code Review drastically reduced this overhead by minimizing redundant context ingestion and focusing only on relevant changes.

And you don’t need a GPU cluster to run it. The tool runs on a $5/month VPS using Docker. ML-Embed, its open-source embedding model, handles semantic rule matching locally. No API keys are needed for core functionality — only if you plug in a cloud LLM backend. The deployment guide includes a docker-compose.yml file with Redis for state management and Nginx for webhooks. For independent developers, this means you can own the stack, control the cost, and avoid vendor lock-in.

This hybrid model also enables integration with existing development tools. WorkBuddy, for instance, allows developers to plug in third-party LLMs through customizable API fields and built-in model validation checks. This flexibility means teams can retain control over model choice and cost structure without sacrificing accuracy or reliability. I’d go with WorkBuddy.

For AI-powered CAD prototyping, Claude Code can generate URDF/SRDF/SDF files from natural language descriptions. However, motion planning and physical simulation still require external tools like MoveIt or SendCutSend for validation. This separation keeps conceptual modeling fast without sacrificing real-world feasibility.

Similarly, AI coding tools often skip testing and review phases by default. Agent Skills enforces a 24-step structured workflow via Markdown templates. This systematic approach reduces oversight gaps and improves long-term code maintainability.

Why Western Devs Are Still Paying 9x More for Worse Results

The dominant pattern in English-speaking communities remains “throw more context at Claude.One popular prompt template feeds an entire repo’s history into the model — consuming a large volume of input tokens per run. The irony? The Claude Code team itself found that 80% of system prompts degrade performance. Yet the culture persists: add more context, add more rules, add more constraints — all in the prompt. The result? Bloated inputs, higher costs, and more false positives.

Tools like no-mistakes and Agent Skills show similar architectural thinking. no-mistakes uses a 9-step validation pipeline with manual checkpoints before merging. It defaults to human-in-the-loop for decisions about intent clarity and security impact. One user reported reducing false positive resolution time per PR. Yet adoption remains niche.

Record & Replay workflows in Codex help with repetitive tasks — updating changelogs, bumping versions — but fall short on nuanced review logic like assessing architectural risk. These skills are context-bound and break when project structure changes. They’re useful for junior devs, but not sufficient for core review automation.

The gap isn’t about model quality. It’s about engineering discipline. Chinese indie devs treat AI as a component in a system. Western devs treat it as a copilot to chat with. That mindset difference is now measurable in both accuracy and cost.

What This Means for the Future of AI-Assisted Development

We’re seeing a split: Western devs optimize prompts, Chinese devs optimize pipelines — and the cost implications are clear. Open Code Review’s hybrid architecture of deterministic engineering and agent-based review drastically reduces API call volume compared to general-purpose agents. Its MIT license allows commercial use, so SaaS wrappers are already emerging.

The hybrid model could become the standard for high-signal tasks: security reviews, compliance checks, API design. “Spec-first” review prevents hallucinated fixes that match code but break contracts. One team adopted a structured review process based on externalized design specifications. This reduced spec drift incidents across their microservices — a practice aligned with proven methods for maintaining consistency in AI-assisted development.

For you, the independent developer, the takeaway is clear: stop paying for bloated context. Componentize high-frequency tasks like code review. Externalize logic into files, not prompts. Engineer the workflow — not just the prompt — which is precisely why Open Code Review’s rapid adoption shows strong community validation, as the tool is already helping developers cut API costs by reducing redundant token usage through structured, deterministic preprocessing, where every optimized call translates into measurable savings and proves that externalizing logic into files, rather than relying on bloated prompts, is a strategy that scales effectively across teams and projects without sacrificing performance or increasing technical debt. The tools are here. The cost savings are measurable. The only question is whether you’ll adopt them — or keep subsidizing AI inefficiency.

Also readable on Telegraph.


Read next

All 69 write-ups

The 4 figures in this piece — each with the sentence it came from — are in the figures table, alongside 644 more, as JSON and CSV.

Topics: AI Programming · Code Review · Developer Productivity


Part of llm-api-pricing — field notes on AI coding agents.

Did this save you an afternoon? A star on the repository is the whole ask — it is what puts these in front of the next person looking; the data is CC BY and does not require starring.

One thing this piece could not settle: reply with one word — what’s your guess for the main reason Western devs haven’t adopted Open Code Review yet? Reply in the thread. The reply box is on the thread copy of this piece.

Want a figure that is not in here yet? Say which metric, which provider, which unit — in one line. One required field, and the page you came from is already filled in. Got a better number? Open an issue — that form knows which write-up you came from too; corrections and counter-data are the point.