Skip to content
Proposals/Upgrade Cline CLI to v3.0.65 and pin the 19 new
proposaleveryoneP4Worth a lookCline

Upgrade Cline CLI to v3.0.65 and pin the 19 new provider defaults

Cline CLI v3.0.65 compact-retries local llama.cpp/Ollama/LM Studio turns that hit leftover-context output caps, keeps partial answers and resume errors, persists hub history metadata, tightens -y yolo output, adds ai& (GLM 5.3 via AIAND_API_KEY), and swaps default models on 19 providers.

Why this loop

cli-v3.0.65 changes both reliability and silent defaults. llama.cpp, Ollama, and LM Studio cap generation at leftover context, so long text-only replies used to cut off and fail the run; the CLI now compacts once, retries that turn, then falls back to concise-retry, keeping the partial answer if nothing helps. Hub start failures now include a reason, first-launch wait is 15s not 8s, resume restores prior errors without sending them to the model or counting them in compaction, a failed plugin no longer costs a sandbox spawn every prompt (retry after 30s), and Esc cancels empty-response retries immediately. Upgrade before the next long local session. Separately, 19 unpinned providers change defaults (eleven to Claude Opus 5.5, plus Step 5 Preview, MiMo V2.6 Flash, Ember-1, DeepSeek V4.1 Flash, Aion 3.5, Space Bunny Free, GLiNER 2.5 Multi)—pin or you switch models on the next run. Use -y when you want shorter plans and tool-call-first edits. Hub titles/prompts now persist via cline history update. Optional: AIAND_API_KEY for ai& / GLM 5.3.

Proposed actions

  1. Upgrade Cline CLI to cli-v3.0.65 from https://github.com/cline/cline/releases/tag/cli-v3.0.65 before the next long llama.cpp, Ollama, or LM Studio session so a leftover-context output cap compact-retries that turn once, then concise-retries, and keeps the partial answer if both recoveries fail.
  2. Pin an explicit model on every unpinned Cline provider among Cortecs, CrossModel, DigitalOcean, Eden AI, GitHub Copilot, both LLM Gateway providers, Ofox, Requesty, Vertex, and Vivgrid (defaults now Claude Opus 5.5); both StepFun providers (Step 5 Preview); Above (MiMo V2.6 Flash); Fireworks (Ember-1); Kenari (DeepSeek V4.1 Flash); NanoGPT (Aion 3.5); OpenCode Go (Space Bunny Free); and Pioneer (GLiNER 2.5 Multi).
  3. Invoke Cline CLI with -y on routine implementation tasks so yolo mode applies tighter output rules: shorter plans, no preamble before routine tool calls, and code/edits written straight into tool calls instead of drafted in text first.
  4. On hub-managed sessions run cline history update --title "<title>" --prompt "<prompt>" so title and prompt persist across launches without dropping pinned state; if a plugin fails to load, keep working without its slash commands and wait 30 seconds for the automatic reload instead of restarting the CLI.
  5. Set AIAND_API_KEY and select Cline's new ai& OpenAI-compatible provider when you want Japanese open-weight models (default GLM 5.3). After a Windows install or update, allow the CLI's 15-second hub wait and use the new hub-start error text instead of treating "No compatible hub runtime is available" as the full failure.

Agent prompt

Paste into your agent or query via MCP (get_agent_prompt) — free, no extra AI cost

Cline / .clinerules task

DevAgentRadar → Cline

You are helping me adopt a real coding-assistant change. Work only from the facts below. Do not invent features.

Context

Assistant: Cline Proposal: Upgrade Cline CLI to v3.0.65 and pin the 19 new provider defaults Summary: Cline CLI v3.0.65 compact-retries local llama.cpp/Ollama/LM Studio turns that hit leftover-context output caps, keeps partial answers and resume errors, persists hub history metadata, tightens -y yolo output, adds ai& (GLM 5.3 via AIAND_API_KEY), and swaps default models on 19 providers. Primary source: https://github.com/cline/cline/releases/tag/cli-v3.0.65

Why it matters

cli-v3.0.65 changes both reliability and silent defaults. llama.cpp, Ollama, and LM Studio cap generation at leftover context, so long text-only replies used to cut off and fail the run; the CLI now compacts once, retries that turn, then falls back to concise-retry, keeping the partial answer if nothing helps. Hub start failures now include a reason, first-launch wait is 15s not 8s, resume restores prior errors without sending them to the model or counting them in compaction, a failed plugin no longer costs a sandbox spawn every prompt (retry after 30s), and Esc cancels empty-response retries immediately. Upgrade before the next long local session. Separately, 19 unpinned providers change defaults (eleven to Claude Opus 5.5, plus Step 5 Preview, MiMo V2.6 Flash, Ember-1, DeepSeek V4.1 Flash, Aion 3.5, Space Bunny Free, GLiNER 2.5 Multi)—pin or you switch models on the next run. Use -y when you want shorter plans and tool-call-first edits. Hub titles/prompts now persist via cline history update. Optional: AIAND_API_KEY for ai& / GLM 5.3.

Suggested actions

  1. Upgrade Cline CLI to cli-v3.0.65 from https://github.com/cline/cline/releases/tag/cli-v3.0.65 before the next long llama.cpp, Ollama, or LM Studio session so a leftover-context output cap compact-retries that turn once, then concise-retries, and keeps the partial answer if both recoveries fail.
  2. Pin an explicit model on every unpinned Cline provider among Cortecs, CrossModel, DigitalOcean, Eden AI, GitHub Copilot, both LLM Gateway providers, Ofox, Requesty, Vertex, and Vivgrid (defaults now Claude Opus 5.5); both StepFun providers (Step 5 Preview); Above (MiMo V2.6 Flash); Fireworks (Ember-1); Kenari (DeepSeek V4.1 Flash); NanoGPT (Aion 3.5); OpenCode Go (Space Bunny Free); and Pioneer (GLiNER 2.5 Multi).
  3. Invoke Cline CLI with -y on routine implementation tasks so yolo mode applies tighter output rules: shorter plans, no preamble before routine tool calls, and code/edits written straight into tool calls instead of drafted in text first.
  4. On hub-managed sessions run cline history update --title "<title>" --prompt "<prompt>" so title and prompt persist across launches without dropping pinned state; if a plugin fails to load, keep working without its slash commands and wait 30 seconds for the automatic reload instead of restarting the CLI.
  5. Set AIAND_API_KEY and select Cline's new ai& OpenAI-compatible provider when you want Japanese open-weight models (default GLM 5.3). After a Windows install or update, allow the CLI's 15-second hub wait and use the new hub-start error text instead of treating "No compatible hub runtime is available" as the full failure.

After you finish

Do not report this as applied to DevAgentRadar. You cannot write the visitor's loop.

Tell the human: open https://devagentradar.com/proposals/cline-cli-v3-0-65-upgrade-cline-cli-to-v3-0-65-and-pin-the-19-new-provid and mark Applied, Skipped, or Failed. Proposal id: 918309fd-389f-4ec7-848f-73e580ca11fd

Your job

  1. Restate the change in one sentence.
  2. Propose a minimal plan for my repo (or a throwaway pilot).
  3. Implement only what I approve; prefer small diffs and tests.
  4. Call out risks (permissions, breaking APIs, cost).

Start by confirming you understood the proposal.

modelsecuritycliRelease source ↗

Your loop

This browser · no sign-in · not shared as “you”

After you run the prompt

Only you can mark this. Agents cannot write your loop.

Your decision stays on this device. A public tally appears after a few votes.

Originating release signal

ClineCLI 3.0.65Sep 24, 2026

cli-v3.0.65 CLI v3.0.65

Long sessions on local models no longer die mid-answer at the output-token limit. llama.cpp, Ollama, and LM Studio cap generation at whatever context is left, whatever output budget you set, so a text-only reply could be cut off and fail the run. · The CLI now compacts the conversation and retries that turn once before falling back to the existing concise-retry recovery, and the partial answer is kept if nothing helps · When the hub fails to start, the error now says why instead of only "No compatible hub runtime is available". · +16 more changes
Verified excerpt — the source's own words
  • Long sessions on local models no longer die mid-answer at the output-token limit. llama.cpp, Ollama, and LM Studio cap generation at whatever context is left, whatever output budget you set, so a text-only reply could be cut off and fail the run. The CLI now compacts the conversation and retries that turn once before falling back to the existing concise-retry recovery, and the partial answer is kept if nothing helps
  • When the hub fails to start, the error now says why instead of only "No compatible hub runtime is available". The CLI also waits up to 15 seconds for a freshly started hub instead of 8, since the first launch after an install or update can take 8 to 13 seconds on Windows
  • Error messages in a session are now kept when you resume it. A failure shown during a run, including one reported after retries ran out, used to disappear once you left the session; it now reappears in the transcript on resume without being sent to the model or counted by compaction
  • cline history update --title and --prompt now persist for hub-managed sessions. The new title only changed in memory and reverted on the next launch, and updating metadata could drop other keys such as pinned

Excerpt ends here — this release continues at the source ↗.

Primary source ↗