If your Claude Code usage has started to feel like a second AWS bill, you’re not alone. Most developers route every single prompt — the one-line typo fix and the full module rewrite — through the same expensive frontier model, because changing that routing meant changing your tools. Together Link is Together AI’s answer to that problem: a free, MIT-licensed CLI, released in beta on October 5, 2026, that runs open models inside the coding agents you already use, with zero workflow changes.
In this guide, you’ll install Together Link, launch Claude Code (or Codex, OpenCode, Pi, or the desktop apps) on open models like Kimi K3 and GLM 5.3, learn how the auto router picks models for you, and see exactly how to track — and cut — your AI coding spend. All facts verified October 8, 2026.
The Problem: You’re Paying Frontier Prices for Routine Prompts
Here’s the expensive habit Together Link is attacking. You open Claude Code, ask it to rename a variable, fix an import, or explain a stack trace — and that request burns tokens at the same rate as a deep architectural refactor. Together AI’s launch post puts numbers on the habit: engineering orgs routinely spend tens of thousands to millions of dollars a month on closed models, with the same premium model handling everything from typo fixes to rewrites.
The counter-argument is that open models have closed much of the gap. Kimi K3 and GLM 5.3 are positioned for hard coding work, while GLM 5.3 Flash and DeepSeek V4.1 Flash handle everyday tasks. Together’s own usage data (reported at launch) shows where the momentum is: as of September 30, 2026, it was serving the largest OpenRouter token share for DeepSeek V4.1 Flash (40.8%), GLM 5.3 Flash (28.2%), and Kimi K3 (23.1%).
The idea behind Together Link is simple: keep the harness, swap the model, and shrink the bill.
What Together Link Actually Does
Together Link connects six coding tools — Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi — to open models hosted on Together AI’s infrastructure. The important detail is what it doesn’t do: it runs no local proxy and no background daemon. Each tool talks directly to Together’s hosted gateway.
That design choice matters for two reasons:
- Your normal config stays untouched. Terminal agents get a temporary, per-launch configuration that is automatically removed when the session ends. Claude Desktop and ChatGPT Desktop use separate, reversible profiles you can switch back from at any time.
- Your existing login and history survive. With Claude Code, you keep your existing settings, memory, and history — Together Link only changes the model route for the session.
If you’ve already tuned Claude Code with custom hooks or an AGENTS.md file, none of that changes. The agent is the same; only the engine under it gets cheaper.
Install Together Link (Two Minutes)
You need macOS or Linux — that’s a beta limitation — plus a Together API key. Sign up at Together AI and grab a key, then run:
curl -fsSL https://link.together.ai/install | bash
The installer adds Bun if you don’t have it and places the commands in ~/.local/bin. Make sure that directory is on your PATH, then confirm the install:
togetherlink --version
Next, wire up your API key. An interactive launch will walk you through togetherlink configure automatically if no key is set, or you can export it directly:
export TOGETHER_API_KEY="your-key-here"
If the underlying agent CLI (say, claude or codex) is missing from your machine, Together Link won’t silently fail — it prints the official install command and docs link for that tool, then exits. The CLI also updates itself automatically, so the beta’s shifting commands land without you reinstalling.
Launch Your First Session
Run togetherlink with no arguments and you get an interactive launcher to pick a tool. Or skip straight to the tool you want:
togetherlink claude # Claude Code
togetherlink codex # Codex CLI
togetherlink opencode # OpenCode
togetherlink pi # Pi
togetherlink chatgpt # ChatGPT Desktop (alpha)
Handy shortcuts exist too: tclaude, tcodex, topencode, tpi, and tchatgpt. Everything after the tool name is passed through, so your normal flags still work:
togetherlink claude --continue
togetherlink codex exec "run the test suite and fix failures"
For scripts and CI, headless mode works the same way — just close stdin so the process doesn’t hang waiting for input:
togetherlink claude -p "summarise this diff" --output-format json < /dev/null
How the Auto Router Decides
Every session defaults to a virtual auto model. The router classifies each request: straightforward tasks go to cheap, fast open models, and harder problems get frontier capability. Routing happens per request, which keeps prompt caching working — a detail that would otherwise silently inflate your costs.
The exact routing depends on your keys:
- With an Anthropic API key configured: the router can escalate difficult requests to Claude Opus 5.5 — but only inside Claude Code and Claude Desktop, and those calls are billed to your Anthropic account.
- Without one: the router stays entirely on Together-hosted open models.
- Codex, OpenCode, Pi, and ChatGPT Desktop always stay on Together models, regardless.
Inside Claude Code, the familiar /model menu maps tiers to open models. The current lineup (check with togetherlink models; the beta list can change):
| Claude Code tier | Open model routed | Model ID |
|---|---|---|
| Opus | Kimi K3 | moonshotai/Kimi-K3 |
| Fable | GLM 5.3 | zai-org/GLM-5.3 |
| Sonnet | DeepSeek V4.1 Flash | deepseek-ai/DeepSeek-V4.1-Flash |
| Haiku | MiniMax M3 | MiniMaxAI/MiniMax-M3 |
The four headline models each ship with 1M-token context, which is the real reason this works for agent loops: you’re not losing the long context window that makes agentic coding useful.
Pin One Model for a Whole Session
Auto routing is a sensible default, but sometimes you want deterministic behaviour — a benchmark run, a like-for-like cost comparison, or just a model you trust for a gnarly refactor. Pin a model by placing the --main flag before the tool name:
togetherlink --main zai-org/GLM-5.3 claude
togetherlink --main MiniMaxAI/MiniMax-M3 opencode
togetherlink --main moonshotai/Kimi-K3 claude
Two gotchas worth knowing:
- The shortcuts like
tclaudeput the tool name first, so they can’t take--main. Use the long form when pinning. - Together Link rejects Claude’s own
--modelflag. Model selection happens through--mainbefore the tool name, or not at all — stay on Auto otherwise.
A Routing Strategy That Actually Saves Money
The router is good, but you’ll save more with deliberate habits. Think in three buckets:
Bucket 1: Chores → Flash models
Renames, import fixes, docstring generation, explaining errors, writing commit messages. These are high-volume and low-risk — the perfect fit for DeepSeek V4.1 Flash or GLM 5.3 Flash. If you’re doing this interactively, leave Auto on; if you want to be strict, pin a flash model for the session and only escalate when the model gets stuck.
Bucket 2: Standard feature work → GLM 5.3
New endpoints, component builds, test coverage, migrations — work that needs real reasoning but not frontier brilliance. GLM 5.3 is Together’s mid-tier pick for exactly this: strong enough to write production-shaped code, cheap enough to run all day.
Bucket 3: Hard problems → Kimi K3 or Opus 5.5
Architecture decisions, tricky refactors, debugging heisenbugs across a large codebase. This is where you spend the premium tokens — Kimi K3 for frontier-class open coding, or let the router escalate to Opus 5.5 if you’ve connected an Anthropic key. The savings come from keeping buckets 1 and 2 off the premium models, not from starving the hard work.
The one-line version: let cheap models take the first swing at routine work, and reserve the expensive ones for work that actually needs them.
See What You’re Spending
Together Link’s killer feature might be its receipts. Every session prints token and dollar totals on exit, so you see the cost of what you just did — not a surprise at month’s end. Inside Claude Code, the status line shows your estimated spend next to what the equivalent Opus session would have cost.
For a weekly view, run:
togetherlink usage --last 7d
That shows gateway-tracked spend across all your sessions. Run it for a week with Auto on, then compare against your previous Claude Code bills. Together claims savings of over 50% — and 50–80% versus all-Opus 5.5 sessions — but treat vendor claims as a hypothesis and verify with your own receipts. That’s the whole point of the built-in reporting.
What the Models Cost Right Now
Here are the launch-week rates listed on Together’s product page (per million tokens; reported October 5, 2026 — confirm live rates at together.ai/pricing before budgeting, since beta pricing moves):
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Kimi K3 | $3.00 | $15.00 |
| GLM 5.3 | $1.40 | $4.40 |
| DeepSeek V4.1 Flash | $0.30 | $1.20 |
| MiniMax M3 | $0.30 | $1.20 |
Billing runs on your existing Together key — pay-as-you-go or credit packs. And a bonus you might not expect: Claude Code, Claude Desktop, and ChatGPT Desktop sessions can also generate and edit images through the CLI:
togetherlink image generate --prompt "a red circle on a white background" --out logo.png
togetherlink image edit --image logo.png --prompt "make the circle blue"
Switching Back Is One Command
Trying a beta shouldn’t feel like a commitment. Because Together Link never touches your normal configs — terminal sessions use temporary per-launch setups, and the desktop apps get separate reversible profiles — going back is trivial:
togetherlink claude-desktop off
togetherlink chatgpt off
togetherlink restore
To fully remove the desktop profiles (it asks for confirmation; add --yes to skip):
togetherlink claude-desktop reset
togetherlink chatgpt reset --yes
One privacy note: Together Link collects anonymous usage analytics by default. If you’d rather opt out, export this before launching:
export TOGETHERLINK_TELEMETRY_DISABLED=1
Limitations Worth Knowing Before You Commit
Honest assessment, because this is a beta:
- macOS and Linux only. No Windows support during beta.
- Commands and routing can change. Together warns that the model list, routing behaviour, and CLI commands may shift. The auto-updater means you’ll get changes without reinstalling — but pin your workflows loosely.
- The Opus path needs an Anthropic key. Without one, you’re entirely on Together-hosted models, which is fine for most work but worth knowing before you expect Opus-tier escalations.
- You’re still paying Together. “Free CLI” means the tool itself is free and MIT-licensed — inference is billed through your Together API key. Run the numbers with
togetherlink usageinstead of assuming it’s cheaper than whatever you’re doing now.
Together Link vs the Alternatives
Model-routing shims aren’t new. Here’s how the field looks:
- OpenRouter: the big catalogue — any model, any agent, via env vars. Maximum choice, but you’re assembling the routing logic yourself.
- Claude Code Router: MIT-licensed, rule-based routing across 10 agents and any provider you configure. More control, but you run and maintain a local gateway (port 3456).
- Ollama launch: free local models in Claude Code with no cloud bill at all — at the cost of your machine’s GPU and smaller local models.
- Together Link: the curated one-command path: no local proxy, per-session cost receipts, and a model lineup pre-picked for coding quality.
If you want maximum provider control, Claude Code Router is the deeper tool. If you want the lowest-friction way to stop overpaying this week, Together Link’s edge is that there’s almost nothing to configure.
Should You Install It?
My take: if you run Claude Code or Codex daily on a paid tier and your usage dashboard has ever made you wince, yes — install it this week, run Auto for a few days, and check togetherlink usage --last 7d against your old bills. The reversible setup means the experiment costs you nothing but an hour. If you barely touch coding agents, or your employer already negotiates your model spend, the savings won’t move your needle and the beta churn isn’t worth it.
The bigger story is the direction this points: model routing is becoming infrastructure, not a purchasing decision. The harness you like and the model you pay for are decoupling — and tools like Together Link are the first one-command version of that future.
Further Reading & References
- Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode — MarkTechPost launch coverage (Oct 5, 2026): install details, auto router, pricing, and the alternatives comparison table.
- Together Link: Free CLI to Swap Models in Claude — AIDailyPost (Oct 6, updated Oct 7, 2026): supported tools, MIT licensing, and why the model-swap matters.
- Together Link Brings Open Models Into Claude Code, Codex, and OpenCode — Without Changing Your Workflow — Tech To Heart: the “model-layer upgrade, not workflow rewrite” angle.
- nutlope/togetherlink on GitHub — the public repo: CLI command reference, model lineup table,
--mainpinning rules, usage tracking, and changelog. - Meet Together Link: Run Open Models Inside Claude & Codex — TechnoSports (Oct 6, 2026): supported-tool rundown and the cost-cutting pitch.
- The Terminal Wars: Comparing the Major CLI AI Coding Agents — command cheat sheet and model comparison across Claude Code, Codex CLI, OpenCode, Pi, and more (Oct 4, 2026).
Facts verified October 8, 2026. Together Link is in beta — commands, routing, and the model lineup may change; run togetherlink models and check the official docs for the current state before relying on any detail here.



