AI Coding Field Notes

Coding agent pricing comparison: cheapest LLM API for coding, cost per million tokens, from the OpenRouter price list as JSON and CSV. An AI model token price comparison repriced at a real agent mix — 95.6% cached input pricing, 6.5x off list: prompt caching price, cache read price, cache hit share, long context pricing, DeepSeek peak/off-peak.

View the Project on GitHub xyzs996/llm-api-pricing

How a Multi-Agent AI System Made $7,600 in 7 Days for Under $100 in API Costs

How a Multi-Agent AI System Made $7,600 in 7 Days for Under $100 in API Costs

Written with AI assistance. Figures without a traceable source were cut before publishing.

The same piece is posted as a thread on GitHub — that copy has a reply box under it, and this one does not.

An AI crypto trading bot reportedly made $7,600 in a week.” That line gets clicks. It also gets rolled eyes. The real story isn’t the profit — it’s the cost: under $100 in API fees. One developer pulled this off using Claude Opus 5, a fully automated multi-agent system, and a cycle that ran every 4 hours. The net gain was around 55,000 RMB (~$7,600), sustained for 7 days. No moonshot luck. No hidden VC funding. Just structured roles, tight loops, and a design that treated resilience as code.

This wasn’t a set-and-forget script. It didn’t run on magic. But it did prove something quietly radical: with the right architecture, you can run powerful AI systems at retail prices. And if you’re trying to monetize AI without burning cash, that changes everything.


Why the “Build Once, Profit Forever” Myth Fails in AI Trading

The dream goes like this: you train a model, wire it to an exchange, and let compound gains roll in while you sleep. Sounds clean. Feels passive. Sells well on Twitter.

It fails in practice because markets don’t sleep — they shift. A strategy tuned for bullish momentum collapses the second volatility spikes or sideways ranges settle in. One developer saw severe losses within two weeks by relying on static prompts without sentiment feedback. The model kept executing trades based on outdated logic, blind to the tone shift in social chatter and on-chain flow.

The illusion of automation hides the maintenance burden. Real systems need prompt tuning, data validation, and risk caps updated weekly — sometimes daily. The bot that generated returns didn’t escape this. It re-ran its strategy calibration every 4 hours, pulling fresh on-chain metrics and social sentiment. Each cycle included verification steps to filter out hallucinated signals before any trade was issued.

And yes, a human reviewed the output on 3 of the 7 days — not to pick trades, but to adjust limits and stop-loss thresholds. Unchecked, the system nearly over-leveraged during a pump-and-dump cycle that looked, for a moment, like a breakout. The override saved the week. I’d take the human override on day 3.

Here’s the hard truth: raw model strength doesn’t guarantee returns. Opus 5 wasn’t chosen for speed or flash reasoning. It was picked because it held context across multi-step agent interactions without degrading. Simpler models failed when the Market Analyst passed data to the Decision Agent — the context slipped, the triggers misfired.

The system used structured JSON for inter-agent communication. That eliminated the ambiguity inherent in free-text handoffs, which is critical when AI is making trading decisions. Unclear instructions introduce risk—and without strict control, risk quickly compounds into real financial loss.

AI trading systems also fail when they can’t integrate real-time actions beyond data processing. For example, an AI that only analyzes sentiment without being able to trigger an inventory check or adjust a live bid risks making decisions disconnected from operational reality. Respond.io’s AI Agent shows how connecting to external systems—like inventory or booking platforms—can close this gap. The system acts on real data in real time. This integration capability is the difference between insight and action.

Similarly, treating AI as a tool rather than a results engine often leads to poor user retention, which was evident in one case where a SaaS-style AI product saw only 20% month-on-month renewal until it shifted focus from selling access to delivering guaranteed outcomes for high-revenue clients, a transition that ultimately revealed how value perception changes when success is contractually assured rather than left to user interpretation. After repositioning its offering around results, renewal rates climbed to over 60%. This shift—from tools to results—reduced user learning costs and increased trust, proving that outcome-based promises resonate more deeply with users who want certainty, not complexity.

Where the Real Leverage Lies: Multi-Agent Specialization

It used three specialized roles: Market Analyst, Sentiment Analyst, and Decision Agent. Each had a narrow job, a fixed input schema, and clear failure boundaries.

Every 4 hours:

  1. Market Analyst pulled price action, volume, and on-chain metrics from Dune Analytics.
  2. Sentiment Analyst scraped Twitter, Reddit, and Telegram using keyword clusters and emotional valence scoring, which fed into the decision-making pipeline where the Decision Agent synthesized both reports, checked them against risk rules, and issued trade commands via Binance API with pre-set position sizing, ensuring alignment across data sources and execution protocols while maintaining strict adherence to predefined strategies [as outlined in the system architecture]. I ship apps; my money’s on Dune Analytics.

No agent saw everything. No agent made assumptions. The architecture forced modularity.

This isn’t just a cute design pattern. It’s how enterprise AI systems avoid cognitive overload. Think of Miora’s multi-agent framework: you assign Standard, Pro, or Max tier agents based on task complexity. You don’t burn Opus-level cost on data parsing — you save it for synthesis and decisions.

The same principle applies here. Separating research from judgment improves reliability and allows for better cost optimization. In this setup, only the Decision Agent used Opus 5, while the other roles relied on more cost-effective models with carefully designed prompts. This strategic allocation reduced expenses, keeping the overall LLM costs low throughout the testing period.

Parallel processing unlocked more than cost savings. It enabled safer iteration. A failed sentiment run? Discard it. The market analysis still updated. Each agent kept memory logs, so failed trades could be audited, and prompts refined post-cycle.

And when disagreement happened — say, the Market Analyst flagged a breakout but the Sentiment Analyst saw FUD piling up — the system used a voting mechanism. No consensus? Trade delayed. Human notified. No blind execution.

You don’t need exotic tech to copy this. You need discipline: define roles, enforce boundaries, and let agents fail cheaply.

This agent specialization model extends beyond trading. In content-driven businesses, similar role separation can drive efficiency: one agent identifies trending topics using APIs like TikTok’s search, another drafts material based on structured prompts, and a third evaluates performance and suggests revisions, mirroring the workflow of a real editorial team.

Another example is in virtual product maintenance, where developers split routine updates into three AI roles: an Information Collector fetching user feedback and market data, a Note Producer generating update logs and changelogs, and a Data Reviewer analyzing engagement metrics to guide next steps. This automation cut manual work from hours to minutes.

Even enterprise AI adoption follows this path—not by launching full-scale systems, but by encouraging employees to use personal agents in daily, low-risk tasks like meeting note summarization or email triage—which in turn generate individual use cases that produce real usage patterns, which organizations then analyze to build scalable, high-impact AI workflows, where the accumulated data from these personal applications informs broader strategic decisions while ensuring alignment with operational realities and enabling gradual, evidence-based scaling across teams and functions.

The Hidden Infrastructure: Aggregation Platforms That Make This Possible

Let’s be honest: most indie devs can’t manage 3+ LLM endpoints, 2+ data APIs, and a patchwork of billing dashboards. That’s why middleware like Novo DasTalk matters.

In this system, Novo DasTalk wasn’t just convenient. It was necessary. It let the developer monitor all agent costs in one dashboard. A budget alert was set: halt operations if daily spend exceeded a predefined threshold. That kind of guardrail doesn’t exist when you’re juggling OpenAI, Anthropic, and Mistral tabs.

Security wasn’t an afterthought. All trade decisions flowed through a private VPC. Agent-to-agent messaging was encrypted. No API keys lived in prompts. Access used short-lived tokens. And data isolation was enforced: sentiment feeds couldn’t leak into the market model and bias its signals.

These backend capabilities turned a fragile script into something production-grade. Unified logging meant debugging a failed trade took minutes, not hours. Versioned agent configurations allowed rollback to a prior state with one click.

The five core functions — model aggregation, interface routing, data handling, billing, and access control — cut dev ops time by about 60%. That’s hours you don’t spend firefighting. Hours you can spend on strategy.

If you’re building AI systems that touch money, you can’t skip this layer. It’s not overhead. It’s insulation.


How to Replicate This — Without Blowing Up Your Account

You don’t need to be a quant to try this. But you do need to start safely.

First: simulate before you risk. The developer ran multiple iterations on Binance Testnet, filtering out several versions due to false-positive bias, which led to the refinement of the final model that performed well in a series of simulated trades and was further stress-tested against historical market shocks — such as the FTX collapse — to evaluate its behavior under extreme conditions where resilience and accuracy were critical for deployment decisions.

Second: build hard circuit breakers. This system had three:

And here’s the critical part: after any pause, a human had to re-enable the bot. No silent restarts. Every trade was logged to a Google Sheet — timestamp, rationale, outcome — for weekly review.

Third: expect to iterate. The developer used Obsidian with claude-obsidian to maintain a research loop on new indicators and sentiment sources, continuously refining the trading strategy through weekly updates. Each iteration involved reviewing losing trades and tightening prompt constraints to improve performance over time.

This system worked because it was narrow, monitored, and built to fail safely. You can replicate the structure. But don’t skip the boring parts. They’re why it lasted 7 days — not 7 hours.

Also readable on Telegraph.


Read next

All 70 write-ups

The 13 figures in this piece — each with the sentence it came from — are in the figures table, alongside 648 more, as JSON and CSV.

Topics: AI


Part of llm-api-pricing — field notes on AI coding agents.

Did this save you an afternoon? A star on the repository is the whole ask — it is what puts these in front of the next person looking; the data is CC BY and does not require starring.

One thing this piece could not settle: reply with one word — what was the name of the analytics platform you’d trust most for on-chain data, based on the author’s hint? The reply box is on the thread copy of this piece.

Want a figure that is not in here yet? Say which metric, which provider, which unit — in one line. One required field, and the page you came from is already filled in. Got a better number? Open an issue — that form knows which write-up you came from too; corrections and counter-data are the point.