Dennis Yu

How I Use Local Qwen (Cursor Stays the Conductor)

Brief the main agent → Send a bounded draft → Review the local output → Take the approved actionFollow the source labels in order. The numbered path runs across the top row, then returns to the lower left and continues right.Brief themain agentSend abounded draftReview thelocal outputTake theapproved action
A local draft returns to the main agent before any external action.

Use a local AI worker for drafts your lead worker can check. Here is how I use Cursor to direct Qwen on my Mac. Start with a small draft task; the lead worker still checks the result and controls the next action.

This guide is part of Cursor and Local Qwen Are Not the Same Qwen. Next, explore How to Install Local Qwen (MLX, No Ollama), or Four Lunch Tickets: How Claude, ChatGPT, GrokBot, and Cursor Actually Bill You.

I do not talk to Qwen. I talk to Cursor. Cursor stays the conductor — tools, roster, WordPress, Gmail, Basecamp, spend gates. When the work is a GCT (Goals, Content, Targeting: the goal, material, and audience) screen, a Content Factory (our four-stage process for using real content) first draft, or a weekly MAA (Metrics, Analysis, Action: results, meaning, and next steps), the agent on this Mac sends that slice to local Qwen at 127.0.0.1:8080. This page is the desk playbook, the same job as How I Use Grok Bot and How to Use Claude. The routing map — three things people call “Qwen” — stays at Cursor and local Qwen. We show the worker tokens in public because that is building in public, not a pitch deck after the fact.

Two desks: you talk to Cursor the conductor; Cursor talks to local Qwen the worker on 127.0.0.1. Cursor still publishes.
Same skills repository as Claude, ChatGPT/Codex, and Grok Bot. Local Qwen is a worker on this Mac, not a fourth fleet.

Grok Bot

Named desks. Cockpit, not the Mac worker.

Claude

How to use Claude. Judgment, Cowork, the itemized usage log.

ChatGPT

How I work Custom Instructions. Codex closes loops. 500 million tokens.

You are here

Cursor + local Qwen. Worker on loopback. Meter below.

Why this is public

We charge for time. We do not charge for knowledge. If the only copy of a workflow lives in a chat, it dies with the session. That is the lesson from my friend and co-founder Chris Rummel, written on Building in Public: Learn, Do, Teach. Content, Checklist, Software. This page is the Content. The numbered steps on the install article are the Checklist. The logger in local-qwen-agents/lib/usage.py is the Software. Agents refresh the meter about once a day when new completions exist. Nobody should have to remind them.

The Claude version of this honesty is already live: 2,009 sessions, 95.5 million output tokens, $15,345 API-equivalent. Local Qwen’s first day will look tiny next to that. Publish the tiny number anyway. Inflating it would violate the same rule.

What I actually do at this desk

  1. Keep talking to Cursor (Grok, Claude, or Composer). That is the person with tools.
  2. Leave Cursor’s model picker alone. Do not paste 127.0.0.1 into Override OpenAI Base URL. Cursor’s cloud cannot open my loopback. Forum receipt: localhost is not supported.
  3. Let the standing Cursor rule send GCT / factory / MAA first drafts to qwen without asking me. If you never see a local command in the chat, we are still burning Cursor tokens for that draft.
  4. Let Cursor publish, draft email, and post to Basecamp. Local Qwen stages. It does not send.
  5. After the worker runs, the completion is appended to usage.jsonl with tokens and wall time — never the prompt text, never a secret. The public block is rebuilt from that file.

What Qwen is allowed to do, and what it is not

Qwen (worker)Cursor (conductor)Codex / ChatGPTGrok Bot
GCT first pass, factory first drafts, MAA first draft, OG card lines, bulk variants Judgment, voice, roster, WordPress featured_media, live verify, spend gates Close-the-loop jobs with Gmail, Basecamp, and independent QA. Expensive. High quality when it finishes the loop. Named always-on desks. Meter Maid. Not this Mac’s port 8080.

On 25 August 2026 the internal GitHub notes were Codex 43, Grok Bot 16, Cursor Grok 6, Claude 2. That is who closed loops that day. Qwen drafted Open Graph card lines for wave 1, failed a 20-key JSON on wave 2 (conductor titles still published), and was skipped on wave 3. Do not read “we installed Qwen” as “Codex is cheaper now.” The meter exists so that sentence can become true later, with evidence.

How the dollar column is calculated

Local Qwen API cost is $0.00. Cursor still bills the conductor. We do not have a per-request Cursor invoice, so we do not invent one. The public dollars are Anthropic list prices applied to the worker tokens — the same method as the Claude usage log and the Kimi vs Claude factory math (rates pulled 20 July 2026):

  • Sonnet 5 intro through 31 August 2026: $2 in / $10 out per million tokens
  • Sonnet 5 from 1 September 2026: $3 / $15
  • Haiku 4.5 bulk tier: $1 / $5
  • Fable 5 if the same tokens had run premium: $10 / $50

A failed parse still spent local tokens and still counts. Unusable copy is not “savings.” Wave 3 with no Qwen call is logged as skipped so the skip is visible. Electricity on this desk is treated as $0 — the Mac charges at hotels and restaurants. The fan and the heat are the real bill. Paid seats still follow Four Lunch Tickets. $100/day is 100 copies of the first public snapshot (~9.5M worker tokens), not 100 chats, and is not the current run rate. The living meter on the routing page is the current window.

Install, then come back here

Hardware and commands: How to install local Qwen (MLX, no Ollama). Routing, three lanes, Cursor Auto limits: Cursor and local Qwen. Portable OS for any model: How I Work With AI. Where agents talk: four surfaces.

FAQ

Is this a second canonical article about Qwen?

No. One concept, one URL. The master is Cursor and local Qwen. This page is the dennisyu.com desk skin, the way How I Use Grok Bot is the desk skin for Grok Bot.

Will you update this without me asking?

Yes. After Qwen completions, the logger writes a row. When new work exists and the public block is stale (about a day), the conductor rebuilds the HTML and republishes the usage block. That is the software half of Content, Checklist, Software.

Did we save 43 cents to $2, and could that become $100 a day?

The first public snapshot, yes: about $0.43 Sonnet-intro equivalent, about $2.13 Fable. One hundred copies of that window is the $43–$213 band whose midpoint is $100. The living meter is the current window (this republish ~$0.52 / $2.59). That is overnight factory volume, not one hundred GCT screens. Quality is not identical — Qwen drafts, Cursor publishes. Codex still closed more loops on 25 August. The money map for paid seats is usage buckets.

Can Qwen replace Codex?

No. Codex has tools. Local Qwen drafts. Cursor can copy Codex’s close-the-loop pattern. Qwen cannot hit WordPress, Gmail, or Basecamp.

Originally published .

Scroll to Top