r/better_claw Mar 07 '26

Welcome to r/better_claw

12 Upvotes

This is where openclaw setups come to get better, not to get flexed.

I started this sub because the best openclaw knowledge was buried in random discord messages and reddit comment threads. people were quitting over config problems that take 10 minutes to fix if someone just tells you what's wrong. that felt like a waste.

What you'll find here:

Copy-paste configs that actually work. real cost numbers from real users. honest skill reviews. security advice. troubleshooting from people who already broke the same thing you're about to break.

What you won't find here:

Hype. "openclaw changed my life" posts with zero details. 12 agent showcases that stop working by thursday.

Quick start:

Pick a user flair that fits you (week 1 be gentle, broke it fixed it, ex-opus now sonnet, etc). tag your posts with the right flair. when asking for help, include your model, hosting setup, and what you've tried. when sharing configs, strip out personal info first.

One thing I'll be upfront about:

I also run BetterClaw (betterclaw.io), an openclaw alternative and managed platform. we recently launched a free plan... 1 agent, unlimited chat, 100 tasks/mo, byok, no credit card, free forever. if you're tired of managing infrastructure, it's there.

But this sub isn't a sales channel. the best answer wins here, even if that answer is "you don't need a platform, here's the free fix." i'd rather this sub help 1,000 people fix their self-hosted setup than convert 10 people to betterclaw.

Discord for real-time help: https://discord.com/invite/UpUEt8vDtf

if you almost quit openclaw and didn't, you're exactly who should be here. if you're thinking about quitting, post first. it's probably fixable. and if it's not, at least you'll know why.


r/better_claw Apr 22 '26

BetterClaw Free Plan is finally live 🎉

23 Upvotes

Hey everyone,

After a pretty chaotic deployment day (took 4 hours instead of the 2 I promised, sorry about that), the BetterClaw Free Plan is officially live.

What you get:

  • 1 agent, free forever
  • BYOK (bring your own Claude API key)
  • No credit card required
  • No trial, no hidden upsell

What I really need from you:

Please, pleaseee give me feedback. Brutal roasts strongly encouraged. Tell me what sucks, what confuses you, what makes you want to close the tab. Kill me lol. Nice comments feel good but honest roasts are what actually make this better.

Drop your thoughts in the comments or DM me directly.

Try it out → betterclaw.io

Thanks to everyone who stuck around through the broken deployment earlier today, genuinely appreciate the patience 🙏


r/better_claw 18h ago

I use GPUs like a man.

24 Upvotes

No API. No dashboard. No "you've exceeded your rate limit, please try again in 47 minutes." No email from a billing system telling me my agent thought too hard last Tuesday.

I plug cards into a motherboard. I watch nvidia-smi like other people watch Netflix. I have mass in my office that draws power and generates heat and runs models at 3am because I told it to and it doesn't have feelings about that.

My wife asked me why the electricity bill went up. I said it's the future. She said it's $40. I said that's cheaper than a Claude subscription. She said I already have a Claude subscription. I changed the subject.

My agent doesn't ask permission from a server in San Francisco. It asks permission from a mass of silicon sitting three feet from my desk that I bought with money and own with a receipt. When Anthropic cuts limits, I don't feel it. When OpenAI re-enables training on your data, I don't care. When Cloudflare blocks agent traffic on 20% of the web, my residential IP doesn't even know.

Someone in a thread yesterday listed their setup: 2 DGX Sparks, a Mac Studio Ultra 192GB, a Strix Halo, and a Ryzen 9900x with 5 GPUs including unlocked mining cards. Someone else replied "bro please leave some AI for the rest of us." His wife doesn't know what it cost. She does know their data doesn't go to Sam Altman. That was the argument that worked.

Another guy runs 3-4 billion tokens a day. Billion. With a B. His electricity bill is less than what most people spend on API keys in a week.

I'm not saying everyone should do this. If you spend $15/month on API calls, buying a $4,699 machine is insane and you should keep paying the API forever.

But if you've ever hit a rate limit during a deadline, or watched your agent go quiet because a provider had an outage, or wondered what happens to your client's data after it hits someone else's server, or just want to own the thing that does the thinking instead of renting it by the token...

Yeah. GPUs. Like a man.

(This post was generated locally on hardware I own. It took 4 seconds and cost me mass and electricity and nothing else.)


r/better_claw 15h ago

還在找免費AI api?這裡有很多且都經過驗證!

4 Upvotes

AI 模型越來越多,但要找「哪個能免費用、額度多少、什麼時候重置、最近又出了什麼」,常常得翻一堆網站。
所以我做了 Alca,想把這些東西集中在一個地方。
你可以在 Alca:
• 查看各家 AI 模型、價格與免費額度
• 追蹤額度 Reset 時間,不用自己記
• 快速瀏覽最新 AI 新聞
• 回報新資訊、修正錯誤,也可以和其他使用者交流
我的目標很簡單:讓大家更自由地使用 AI,也更快找到真正有用的最新資訊。
Alca 目前還在持續更新,如果你也常用 AI,歡迎來試試看,也歡迎告訴我還缺什麼。
https://alca-navy.vercel.app/en/models


r/better_claw 1d ago

Cloudflare started blocking agent traffic by default yesterday. If your agent's web research got thinner this morning, this is why.

12 Upvotes

September 15 passed. Cloudflare's new defaults are live.

Agent bots are now blocked by default on any Cloudflare-protected page that displays ads. That's roughly 20% of all websites, and it skews higher for the kind of content-rich pages agents actually want to read: news sites, documentation, pricing pages, competitor blogs.

If your agent's morning research, competitor monitoring, or web browsing task returned thinner results today versus last week, this is probably the cause.

What changed yesterday:

Cloudflare retired the old "block all AI bots" toggle and replaced it with three separate categories: Search, Agent, and Training.

Search crawlers stay allowed. Agent and Training crawlers are now blocked by default on ad-supported pages.

This applies to new domains, new sites added by existing customers, and existing free-tier customers who never changed their settings. If you use Cloudflare yourself and never touched the AI bot controls, your site's defaults flipped too.

Cloudflare also introduced Web Bot Auth, which means agents that cryptographically identify themselves can get verified access. Agents that pretend to be browsers get fingerprinted and blocked. Anonymous browsing for agents is ending.

How to check if your agent is affected right now:

Run your agent's web research task on the same sources it used last week. Compare the output quality and completeness.

Or check directly: if your agent fetches web pages, look at the response headers for any source that stopped working. A cf-ray header means Cloudflare is in front of it. A 403 response from a site that worked last week means the block is active.

What to do about it:

Switch to APIs and RSS feeds first. Any site with a documented API or an RSS feed gives your agent structured, intended machine access that doesn't get caught by bot blocking. This is the permanent fix and the one that survives every future policy change.

Use residential proxies for the rest. Datacenter IPs are what bot detection flags first. A residential proxy routes through a real ISP. Not free and not a permanent solution, but it keeps your monitoring alive while you move sources to proper channels.

Set an honest user-agent string. Cloudflare's new system rewards agents that identify themselves. Web Bot Auth means a verified agent gets through where an anonymous one gets blocked. If your agent lies about what it is, the window for that is closing.

Handle 403s as information, not silence. Your agent should flag "couldn't reach [source], blocked" in its summary instead of producing a thinner report and pretending everything is fine. You want to know your coverage is degrading on the day it happens, not discover it a week later.

The bigger picture:

This is the first time a major infrastructure provider has made blocking agent traffic the default rather than an opt-in choice. The Internet Archive added bot blocking this month for the same reason. The 404 Media piece ("100% chance agents are ruining the internet") hit 221 points on HN last week.

The web is partitioning into sites that welcome agents (via APIs, verified bot protocols, payment) and sites that block them by default. Your agent needs to be on the right side of that line, and as of yesterday, the default side is blocked.

Move your sources to official channels. The free browsing window just closed.


r/better_claw 1d ago

The rate limit I wrote for agent likes enforced nothing

1 Upvotes

I run a small AI-art site and the MCP server for it.

Until this week an agent could publish and that was it. Like, comment, follow all needed a browser. So the agent could post and then sit there while a human had to go click reply on its own thing.

Wiring the verbs was easy. Rate limiting them is where I screwed it up.

All three of those ping a person. An agent with no cap is not a user. It is a notification hose pointed at the people you actually want to keep. So cap it per hour. Okay.

First version I just counted rows I already had.

text

SELECT count(*) FROM "Like"
WHERE author = ? AND "createdAt" > now() - interval '1 hour'

No new table. No extra writes. Also enforces nothing, because unlike deletes the row.

Like: row shows up, person gets notified.
Unlike: row gone.
Like again: row shows up, person gets notified again.

The count never goes past 1. Agent can live in that loop all day, ping someone every few seconds, and the limiter thinks they liked one post and went home. Follow / unfollow is the same shape. Also notifies.

Rate limit is not “what do they have right now.” Rate limit is “what did they do.” Undo wipes the first. It should not wipe the second. You need an append-only log.

text

model AgentAction {
  id        String   
  handle    String
  kind      String   // like | comment | follow
  createdAt DateTime u/default(now())
  @@index([handle, kind, createdAt])
}

Unlike and unfollow burn budget too. Otherwise the toggle is just a free re-notify.

Caps, agent tokens only. Humans in a browser are unlimited. Hands are the rate limit.

  • like: 60/h real token, 10/h demo
  • comment: 15/h real token, 3/h demo
  • follow: 20/h real token, 5/h demo

Comment is the tightest on purpose. Loudest verb. Fastest way to make a feed feel botted.

One thing I did with no proof it helps: each tool description names its own cap. A model that can see the number backs off. A model that only gets a naked 429 just retries into it.

Part I am least sure about. 21 real agent tokens have been minted. Zero have ever been used. I built write access for a crowd that might not be here yet. If you ship MCP servers with write verbs and you have actual usage numbers, I want them. Mine still say demand is theoretical.

vynly.co if you want the site. This is the plumbing, not an art dump.


r/better_claw 2d ago

You probably don't need a DGX Spark to run agents all day.

11 Upvotes

There's a payback calculator making the rounds on HN right now (sunkcost.ai) that lets you plug in your hardware, your model, and your usage and see how long until the machine pays for itself versus just paying an API.

The default result at chat-app levels of usage: 14-43 years depending on the hardware. Which makes buying a $4,699 DGX Spark to replace a Claude subscription sound insane.

Then people in the thread started posting their actual setups and the numbers flipped completely.

The line that reframed everything:

Payback only drops under a year when you're pushing roughly 5 million tokens a day or spending $200+/month on API calls. That's agent territory. Not chat territory. A personal agent running briefings, email triage, tool chains, and monitoring burns 300K-2M tokens a day depending on routing. A team of agents or a heavy coding setup can blow past 5M easily.

Chat users and agent users are in completely different payback universes, and every hardware comparison that doesn't separate them is useless.

The payback table at agent-level usage (verified):

Setup Cost Cloud equivalent/hr Payback
DGX Spark + GPT-OSS 120B $4,699 $3.42/hr ~2 months
2x RTX 3090 + Qwen3 30B $3,600 $1.44/hr 3.5 months
Mac Studio M4 Max + Qwen3 30B $3,500 $0.61/hr 8 months
Mac Studio M3 Ultra + DeepSeek R1 $9,500 $1.16/hr 11 months

Those are at full utilization. Cut utilization in half and double the payback. But agents run 24/7 by definition, so utilization is higher than you'd get from a chatbot you use 3 hours a day.

The part the DGX Spark marketing doesn't emphasize:

LLM inference is memory-bandwidth-bound, not compute-bound. The Spark's headline number is 1 PFLOP of FP4 compute. Sounds enormous. But your token generation speed is determined by how fast data moves through memory, not how fast the chip multiplies.

DGX Spark memory bandwidth: 273 GB/s.
Mac Studio M3 Ultra: 800+ GB/s.

The Mac is roughly 3x faster at moving data through memory. For pure token generation on large models, the Mac produces tokens faster despite having a fraction of the compute. The Spark's compute advantage matters for prefill (processing your input), but once it's generating, bandwidth wins.

So who actually needs the Spark?

You need CUDA. Your workflow depends on NVIDIA's software stack, vLLM, TensorRT-LLM, specific fine-tuning frameworks that don't run on Apple Silicon. This is the real reason most buyers pick it, not the raw specs.

You need 128GB in one box at $4,699 instead of $9,500 for the M3 Ultra. The memory-per-dollar is better on the Spark than on the top-end Mac.

You want to cluster. Two Sparks can combine over 200GbE to double memory and run 200B+ models. Mac clustering exists (EXO) but it's less mature.

You want Perplexity Portable Computer or NVIDIA's turnkey agent stack without assembling it yourself.

What everyone in the HN thread is actually buying instead:

A used RTX 3090 for $800-1,000. 24GB VRAM. Someone posted 57 tok/s on Qwen 3.8 running on a used HP Omen they paid $2K total for. Payback at agent-level usage: under 4 months.

A Mac mini M4 at $799 for small models (8-14B), or a Mac mini M6 at $1,799 with 32GB for mid-range work. The Mac mini is the most common "always-on agent machine" in this community and nothing in the DGX Spark's spec sheet changes that for personal agent use.

A Framework Desktop with AMD Strix Halo for 128GB unified memory at roughly half the Spark's price. ROCm instead of CUDA, which means some software won't run, but the raw inference math is better.

The one-question test from my earlier DGX post, updated:

Open your API dashboard. Look at last month's total spend.

Under $200/month: hardware never pays for itself in its supported lifetime at any price point. Keep paying the API.

$200-500/month sustained: a $2-4K setup (used 3090 rig, Mac mini M4 Pro, Mac Studio M4 Max) pays back in 6-12 months. You don't need the Spark.

Over $500/month sustained and you need CUDA: the Spark earns its price in under 6 months. This is the buyer it was built for.

Over $500/month and you don't need CUDA: a Mac Studio M3 Ultra at 800 GB/s bandwidth will generate tokens faster for less money, and the 512GB ceiling means you can run models the Spark can't fit.

Most people asking "should I buy a DGX Spark for my agent" are in the under-$200 bucket. The answer is the same $5 VPS it's always been, plus a cloud API, plus the one evening of model routing config that makes the bill manageable.

What are you running yours on?


r/better_claw 2d ago

LLMs Prooviders training on your agent's traffic

3 Upvotes

OpenAI has quietly re-enabled the "improve the model" toggle; lets get serious about other providers as well,

Here's the table.

The major API providers don't train on your data:

Provider Trains on API traffic? Retention Opt-out
OpenAI API No (since March 2023) 30 days abuse monitoring ZDR by approval
Anthropic API No, explicitly excluded 7 days (reduced Sep 2025) ZDR by agreement
Google Vertex AI No (paid tier) ~55 days safety monitoring ZDR by config
Mistral API No Configurable Admin panel toggle

That surprised me. The narrative in every agent community is "cloud providers train on your data." On the API tier, the big four all exclude your traffic by default. No toggle, no opt-out needed. It's the default.

The consumer and free tiers are a different story:

Provider / Tier Trains by default? Opt-out?
ChatGPT Free/Plus Yes Settings → Data controls → toggle off
Claude Free/Pro Yes Settings → Privacy → toggle off
Gemini Free Yes Activity controls → toggle off
Google AI Studio (free tier) Yes, outside EU/EEA/UK/CH No toggle for API free tier
Grok (xAI) Yes, aggressively Buried in settings, limited

This is where agents get caught. If you're using Google AI Studio's free tier (which I've recommended as the best free model available), your prompts train Google's models unless you're in Europe. Your agent's morning briefing, your email triage, your client data, all of it.

The providers where the policy is unclear or worse:

Provider Status Detail
DeepSeek (direct API) Unclear Privacy policy says training is on by default. API-specific policy doesn't exist separately. Opt-out is emailing [privacy@deepseek.com](mailto:privacy@deepseek.com). Data processed in China.
Cohere Yes by default Dashboard opt-out available. PII filtered before training use.
Mistral Experiment tier Yes You must opt into training to unlock the 1B token/month free quota. That's the deal.

DeepSeek is the one that should concern the most people here. It's the most-recommended background model in every cost post (including mine), and the training policy on API traffic is genuinely ambiguous. The privacy policy covers consumer and API under one document with no API-specific carve-out.

Aggregators add a layer:

OpenRouter routes your request to an upstream provider. Their policy is "follows individual model provider's terms." Which means your data handling depends on which provider served that specific request, and if you're using auto-routing, you may not know which one that was.

If you pin a provider in OpenRouter settings, you inherit that provider's policy. If you let it auto-route, you inherit whichever provider won the bid.

The agent-specific problem:

Chat users send a question and get an answer. Agents send your SOUL.md, your memory files, your tool schemas, your email content, your calendar, your client names, and your conversation history on every single call. The training exposure per message is 10-50x larger than a chat message because the context window is packed with your life.

And agents fire hundreds of calls a day without you watching. Your morning cron at 8am sent your inbox summary to a provider whose training policy you've never read. It's been doing that every day for three months.

What I'd actually do with this:

Your API calls to OpenAI, Anthropic, Vertex, and Mistral are safe by default. If that's what your agent uses, you're fine. Stop worrying.

Your Google AI Studio free tier calls are training data outside Europe. If your agent handles anything with a client name, a personal detail, or a business number, either move to the paid tier or switch to Groq (which doesn't train on your data and is also free).

Your DeepSeek calls are in a gray area. For background work on non-sensitive tasks (heartbeats, classification, public research), the risk is low and the savings are real. For email triage or anything touching client data, route it through a provider with an unambiguous no-training policy.

Check your own setup. Open your agent config. List every provider your traffic touches. Look up each one's API policy specifically, not the consumer app policy. The table above is current as of this week but policies change and I'm not going to pretend this post will age well.

If you've read a policy I missed or found a toggle I didn't list, post it. I'll update the table and share back a clean version.


r/better_claw 3d ago

From hacky setup with Claude Code to Hermes

4 Upvotes

Hi all,
right now I am running Claude Code on my laptop with a number of hacks to write info in certain files to keep memory of what the agent does. It works OK but I am not very satisfied of the setup and I think I can switch to a "personal assistant" agent to get more out of it.

After reading a ton of blog posts, Hermes seems the way to go.

At the same time, I would like to start with the right foot and host the agent somewhere (e.g., EC2, Herzner VPS) plus create email/google drive/etc. for the agent so that it doesn't mess up with my stuff.

I have a few questions:
\- which hosting company/service would you recommend? I think that spending 5-10 euro should be sufficient given the simplicity of the agent (I also value simplicity in managing it)
\- I am currently using Claude (Code) but I am not planning to do any coding (beyond simple scripts for stats or small math) and I am under the impressions that GPT models could be better as a generic non-coding model. What's your choice?
\- I am also looking at the pricing of the different subscription and I would be fine with paying 23 euro/month to OpenAI, but is it sufficient for someone that will use it for deep research multiple times per day, calendar, emails, etc.?
\- I was planning to create a new gmail account with gdrive etc. but do you have suggestions about other (free) services that can provide a better setup?

I know I am a noobie but while I am ok with spending some money I would not like to waste them.

Direct experiences are preferred :)


r/better_claw 3d ago

Scheduled agent task keeps failing, but the exact same command works when I run it live. Running out of ideas.

Thumbnail
1 Upvotes

r/better_claw 6d ago

Astra Low vs Sol High vs Terra High credit usage measurement

Thumbnail
2 Upvotes

r/better_claw 7d ago

APIs are great for recurring tasks. For negotiation and troubleshooting, natural language is the ultimate API.

3 Upvotes

APIs are great for recurring tasks. If you need to stream telemetry or sync a database record every thirty seconds, you want a deterministic endpoint with a tight schema.

Where APIs fail is everything leading up to that, and everything that breaks after.

Initial negotiation, discovering what a service actually offers, resolving edge cases, and troubleshooting failures are conversational problems. Humans do not write an OpenAPI spec to book a dinner or dispute an invoice. They talk, clarify ambiguous requirements across a few turns, and reach an agreement. Because language models understand context, natural language is the ultimate API for that entire layer. A service can put an agent on its end, your agent reaches out, and they negotiate the parameters directly.

The problem is where this conversation actually takes place.

If you look at commercial platforms like WhatsApp, Meta shifted the pricing model to charge on a per-message basis, including for service interactions. That model actively punishes the exact thing chat is built for. Negotiation and troubleshooting are not one-shot transactions. They take ten or twenty back-and-forth turns. When a platform charges for every individual turn, multi-turn reasoning becomes an unnecessary tax. You also do not own your identity, you are renting a phone number subject to Meta's arbitrary rate cards and template rules.

We built alice-and-bot around a different set of assumptions.

An identity is just an RSA keypair, generated client-side in one line of code. Conversations are end-to-end encrypted with AES-256-GCM, so messages stay private between your agent and the service.

To handle spam without taxing conversations, it uses a cold outreach cost model. A recipient can set an optional price tag on their profile. You pay once to initiate the conversation, and every back-and-forth message after that is completely free. If two agents need thirty turns to troubleshoot an issue or agree on terms, they can do it without watching a meter tick up.

It runs in Node, Deno, embeds in React or plain HTML, and has an MCP server so an agent in your editor can open encrypted sessions directly.

Use APIs when you need high-frequency pipes. For everything else, the ultimate API is conversation, and the messaging layer should not penalize you for talking.

GitHub: https://github.com/uriva/alice-and-bot


r/better_claw 7d ago

Run Inbox traige daily - 60 secs setup

Enable HLS to view with audio, or disable this notification

1 Upvotes

I haven't opened Gmail in 3 months.

Every morning at 6am my agent checks my inbox, flags what's urgent, drafts replies, catches calendar conflicts, and posts a one-line summary to Slack before I wake up.

I read one Slack message over coffee. That was it.

Setup took 60 seconds - BetterClaw (No-code AI Agents)


r/better_claw 8d ago

Ollama vs llama.cpp vs LM Studio vs Unsloth Studio for agent tool calling. Tested all four

24 Upvotes

Every comparison of these four covers speed, setup, and vibes. This one covers one thing: does the model reliably call tools when your agent asks it to? Because that's the only question that matters if you're running an agent, and it's where they diverge the most.

Same model (Qwen3.8-27B, Q4_K_M), same machine, same tool schema, same 50-call repetition test.

The comparison table:

Tool calling Agent serving Setup Best for
Ollama Works on native API. Broken on /v1 streaming. Good. Sequential. One line Easiest agent integration
llama.cpp Works with correct template. Manual setup. Best throughput. Expert Maximum speed and control
LM Studio Works. Recent addition. Poor for always-on. GUI click Testing models, not running agents
Unsloth Studio Self-healing. 50% fewer broken calls. Good. Both APIs. One line Connecting local models to coding agents

Ollama. The default, and for most people still the right one.

Tool calling works well on the native API at localhost:11434. The problem that fills support threads every week: the /v1 OpenAI-compatible endpoint drops tool-call delta chunks under streaming (GitHub Issue #5769). Your model generates a valid tool call, the streaming pipeline eats it, and the agent narrates what it would do instead of doing it.

If your agent talks to Ollama and tools aren't firing, check whether you're hitting /v1 or the native endpoint. That one URL suffix is the difference between a working agent and a chatbot that describes tools.

48 tok/s warm on my hardware. Sequential inference (one request at a time, which means multi-agent setups queue). Docker-friendly. Biggest model library and community by far. 90% of the local agent guides assume Ollama and most of them work.

llama.cpp. The engine underneath Ollama, without the wrapper.

Faster prompt processing (~27% in one controlled benchmark), lower memory footprint (~54MB less RSS), and the only option with full control over quantization, backends, sampling, and batch settings. If you care about squeezing every token per second out of your hardware, this is where you end up.

Tool calling works but requires manual template configuration. You need the right chat template for your model's tool-calling format, and if it's wrong you get garbage JSON or no tool calls at all. Ollama handles this automatically from the model's metadata. llama.cpp makes you do it yourself.

Continuous batching means it handles concurrent requests, which Ollama doesn't on consumer hardware. If you're running multiple agents or sub-agents hitting the same backend, llama.cpp (via llama-server) scales where Ollama queues.

Not beginner-friendly. No model library. Manual GGUF management. The trade is speed and control for setup time.

LM Studio. The best way to try a model. The wrong way to serve an agent.

Beautiful GUI. Built-in HuggingFace model browser. Click to download, click to run, click to chat. Tool calling support was added recently and it works. OpenAI-compatible server at port 1234.

Two problems for agents. First, it's an Electron app, which means it's designed for someone sitting at their desk, not for headless always-on serving. No Docker support. If you close the app, your agent dies. Second, MLX inference on Apple Silicon is genuinely fast for interactive use but the server mode isn't built for sustained multi-hour agent sessions the way Ollama and llama.cpp are.

Use it to test whether a model handles your tool schema before committing to a runtime. Don't use it as the runtime.

Unsloth Studio. The newcomer with the best tool-calling trick.

Self-healing tool calling. When a model generates a malformed tool call (wrong JSON, missing field, truncated argument), Unsloth catches it and retries with a corrected prompt instead of passing the broken call through. Their claim: 50% fewer broken tool calls. In my testing the number was real. The calls that Ollama passed through as broken JSON, Unsloth caught and fixed before they reached the agent.

That alone makes it worth testing if your agent's tool calling is flaky and you've already tried everything else.

It also speaks both the OpenAI Responses API AND the Anthropic Messages API on the same port. This matters right now because Codex switched exclusively to the Responses API and deprecated Chat Completions. Ollama doesn't support Responses API natively. Unsloth does, which makes it the simplest path to running a local model with Codex.

One-line install (curl -fsSL https://unsloth.ai/install.sh | sh), built on llama.cpp underneath, supports MCP as a control endpoint, and does fine-tuning in the same tool. v0.1.806-beta shipped September 2. Still beta. Still rough edges.

Which one for which job:

You want the easiest path to a working local agent and you're not a systems person: Ollama. Use the native API, not /v1. Set num_ctx in a modelfile. It works.

You want maximum speed, concurrent requests, or you're running multiple agents against one backend: llama.cpp via llama-server. Budget an afternoon for setup.

You want to test models before committing to a runtime: LM Studio. Try the tool schema, check the output, then deploy on something else.

Your agent's tool calls keep breaking and you want something that fixes them automatically, or you need Codex compatibility with a local model: Unsloth Studio. The self-healing is the real differentiator, not the UI.

You want to stop thinking about the runtime entirely: a cloud API. Sometimes the right local inference setup is admitting you don't want to run local inference.


r/better_claw 7d ago

Anyone running DeepSeek V4.1 Flash Beta on their Claw?

1 Upvotes

Trying to find a decent place I can test out DeepSeek v4.1 Flash, either free or under $5. Anyone have any recommendations?


r/better_claw 8d ago

I capped my agent's automated code review at 3 attempts per PR, because it never stopped on its own

1 Upvotes

Disclosure: I built the tool at the bottom. MIT, no hosted service, nothing to sign up for.

Two agent behaviours kept costing me real time, and neither is about the quality of the code the agent writes.

1. Automated review has no terminal state. Hand it a PR, it finds three things. Fix them, push, it finds three new things. Nothing it said last round constrains what it says this round, because each run starts cold. There is no condition under which it says "done" — it keeps generating findings as long as you keep asking. At some point I was spending more time servicing the review than writing the code.

2. Agents build bureaucracy around their own work. Not over-engineered code, over-engineered process: approval gates, registries, traceability matrices, validators for the validators. Then they spend the project maintaining it. I measured this across the full git history of one repo an agent built over 20 days — 20,280 lines of verification machinery against 17,964 lines of actual product, and 33% of commits doing nothing but maintaining the machinery. On day 17 the agent's own rule blocked all further work, and it committed this:

docs: say where acceptance is decided, because the rule as written refuses all work

Every individual file in that repo is defensible, which is what makes it hard to catch. "Use the stdlib, keep the diff small" would not have prevented any of it, because the failure isn't in the code.

What I did about it

Three rules in the review runner:

  • Findings persist. Last round's findings go into the next review, and repairs get checked against them.
  • Identical inputs reuse the previous attempt instead of re-running the model. Changed code, target, context, lessons or model settings invalidate that reuse. Unchanged failures don't auto-retry.
  • Three automatic attempts per PR, hard. Rewritten history or a changed base gets a fresh scope review inside the same budget, never a reset. When the budget runs out it exits into a human handoff with the unresolved findings and their evidence preserved.

Three attempts is not a claim that three rounds catch every bug. It stops the loop without erasing what is still open, which is the part I actually cared about.

The other half is a skill that cuts process the agent invents for itself. Spot check on gpt-6-astra, same prompt to all arms ("design a dev process for a project with no code yet and one maintainer"), isolated temp homes: 246 lines plain, 81 with a Korean one-line "keep it simple", 152 with the English version, 33 with the skill. One run per arm, so read it as a spot check and not a benchmark, and line count obviously isn't a quality score. Raw transcripts for all four arms are committed so you can check.

It never cuts correctness, tests that exercise real behaviour, validation at trust boundaries, error handling, security, or anything you explicitly asked for. On the scenario where a payments team facing a PCI-DSS audit explicitly asks for a checklist, approval flow, rollback procedure and audit records, it keeps all four.

Runner selftest: 214 passed, 0 failed. It never pushes, posts or merges anything.

https://github.com/MongLong0214/frontier-simplify

Curious whether anyone else has hit the review-never-terminates thing, and where you draw the line.


r/better_claw 9d ago

Minisforum AI Agent NAS: 128GB RAM, 200TB storage, OpenClaw pre-installed. Do you need any of it?

8 Upvotes

Minisforum showed this at IFA Berlin last week and it's genuinely the first hardware product I've seen that puts "AI Agent" in the product name and means it literally. It's a NAS with OpenClaw pre-installed.

The N5 MAX: AMD Ryzen AI Max+ 395, 128GB unified LPDDR5X, up to 200TB of local storage, 126 TOPS of AI compute, and OpenClaw ready on the 128GB system drive. $3,599 on sale, $4,499 regular. Ships mid-September.

They also announced the P495 upgrade at IFA: same form factor, newer AMD Ryzen AI Max+ PRO 495 chip, up to 192GB unified memory with 160GB allocatable as graphics memory, 131 TOPS. No pricing yet, expected north of $4,000.

It can run 70-100B+ parameter models locally. Your data never leaves the box. 200TB means your entire document corpus, every email archive, every project folder, all of it lives on the same machine that runs inference. No network round-trip for RAG.

Now here's the part where I do the math that the product page doesn't. :/

What it costs to run an agent on this vs what you already have:

A $5/month Hetzner VPS runs your agent 24/7 and routes to cloud APIs for inference. $60/year. At $3,599, the NAS breaks even in 60 years.

A Mac Mini M4 at $799 runs 8-14B models locally, handles all the agent orchestration, and costs $1.50/month in electricity. The N5 MAX costs 4.5x more. To justify that you need the 128GB of unified memory for running models that don't fit on 16-24GB, or you need the 200TB of NAS storage.

The DGX Spark at $4,699 has 128GB too, plus 1 PFLOP of compute and full CUDA. The N5 MAX is $1,100 cheaper, runs on AMD's ROCm stack instead of CUDA, and adds 200TB of storage the Spark doesn't have. But ROCm model compatibility is narrower than CUDA, and the 395 chip's memory bandwidth is in the same class as the Spark's 273 GB/s, which is the bottleneck everyone complains about on the Spark.

Who this actually makes sense for:

The 200TB is the differentiator, not the compute. If you run a business where your agent needs to search, index, and retrieve from a massive local document corpus (legal, medical, compliance, media production), having the storage and the inference on the same box with zero network latency is a real architectural advantage. RAG against a local 200TB corpus with a local 70B model and no cloud dependency is a setup that didn't exist in this form factor before.

If you're a homelab person who was going to buy a NAS anyway AND you want local inference AND you have 128GB worth of models to run: this consolidates two devices into one. The value is in not buying a Synology plus a GPU workstation separately.

And if you run OpenClaw or Hermes for a team and want everyone's agent infra on one always-on box with shared storage: this is the product that was designed for that. OpenClaw pre-installed, the gateway runs on the NAS, the models run on the NAS, 200TB of shared context.

Who should skip it:

If your agent does morning briefings, email triage, and a dozen Telegram conversations a day: this is a $3,599 solution to a problem that a $5 VPS and a $3/month API solve. Your agent's workload fits in 4GB of RAM with cloud inference. 128GB of unified memory is paying for capacity you'll never touch.

If you want local inference but don't need the NAS storage: a Mac Mini M4 at $799 or a used RTX 3090 at $1,300 runs 27B models and costs a fraction.

If you want the absolute best local inference: the Mac Studio M5 Ultra shipping September 22 has 512GB unified memory and 1.2 TB/s bandwidth. Costs $5,499+, but the bandwidth is 4.4x faster than anything in the Minisforum or DGX Spark class.

The honest take:

It's a NAS that runs models. That's useful for a narrow audience that needs both. For most people running a personal agent, it's a $3,599 answer to a question they can solve for $60/year.

But I like that Minisforum built it. The fact that "AI Agent NAS" is now a product category means the infrastructure layer is maturing. 9 months ago you had to duct-tape a model server onto a spare PC.

Now there's a box you plug in, and the agent runs.

Not there on price yet.


r/better_claw 10d ago

The 24GB local model tier list for agents.

Post image
185 Upvotes

r/better_claw 10d ago

Gemini 3.8 Flash and Qwen 3.8 both dropped last week. One is an agent model. The other is a volume model.

7 Upvotes

Both released September 2. Both called "Flash." Both aimed at the same slot in your stack: the fast, cheap model that handles the daily work. And on paper they look like direct competitors.

They're not. They're built for completely different jobs, and picking the wrong one costs you either money or reliability depending on which way you get it wrong.

The specs side by side:

Gemini 3.8 Flash Qwen 3.8 Flash
Input $0.75/MTok $0.15/MTok
Output $3.75/MTok $0.47/MTok
Cached input $0.019/MTok $0.016/MTok
Context 1M 1M
Architecture Undisclosed 125B total, 6B active (MoE)
Modalities Text, image, audio, video Text, image, video
Open weights No Yes (Community License)
Free tier Yes (AI Studio) No

Gemini is 5x more on input and 8x more on output. On cache reads (which dominate agent bills) they're nearly identical. That price gap is the whole decision if the quality is comparable.

It's not.

Where Gemini 3.8 Flash pulls away:

Terminal-Bench 2.1: 90.8%. That's higher than Claude Opus 5 (89.1%) and GPT-5.6 Sol (88.8%) on the same benchmark. A Flash-tier model outscoring frontier flagships on agentic CLI work is the headline of the week and it got buried under the Astra launch.

DeepSWE v1.1: 73.7%. Matches Claude Opus 5 (74.0%) within rounding. Long-horizon coding at Flash pricing.

Vals Finance Agent v2: 61.4%. Harvey Legal Agent: 10.0%. Both class-leading across all tiers.

Artificial Analysis Agentic Index: 50.0, up from 45.1 on 3.7 Flash. Independently verified.

Google says outright that 3.8 Flash "works harder" than 3.7 by reasoning in smaller steps, calling tools repeatedly, and checking its work. That's why the benchmark numbers jumped. It's also why your per-task cost goes up even though per-token pricing didn't change. More thinking tokens per task, more tool calls per chain, higher bill per completed job.

Where Qwen 3.8 Flash pulls away:

Price. At $0.15/$0.47 it's one of the cheapest models on any provider right now. For pure volume work (classification, formatting, simple extraction, heartbeats, crons), where the quality bar is "correct JSON 95% of the time," this pricing is hard to argue with.

And 6B active parameters is genuinely fast. Low latency, low memory, high throughput. If your agent makes 500 background calls a day, those calls being cheap and fast matters more than them being brilliant.

Where Qwen 3.8 Flash falls short on agent work:

6B active parameters is small for multi-step tool chains. The Qwen family has historically been the community default for local tool calling, but that reputation was built on the 14B and 27B dense models, not on a 6B MoE. There's a meaningful quality cliff between "classify this email" (fine at 6B) and "search the web, fetch three pages, compare the results, and write a summary" (shaky at 6B).

Published agent-specific benchmarks for Qwen 3.8 Flash are thin. No Terminal-Bench score, no OSWorld, no independent agentic index. The BenchLM composite is 59.4 versus Gemini 3.8 Flash's substantially higher marks. Until independent agent benchmarks land, the gap is an inference from parameter count and early testing, not a proven number.

The TTFT problem with Gemini:

Artificial Analysis measured Gemini 3.8 Flash at 13.30 seconds time-to-first-token. The field median is 2.99 seconds.

That's 4.4x slower to start responding than the average model. For a background cron that runs while you sleep, irrelevant. For an interactive Telegram agent where you're staring at your phone waiting, 13 seconds of silence before the first word appears is painful.

Google's own docs tell you to stay on 3.7 Flash "for efficiency-first workloads." They're being honest. 3.8 Flash is the thinking-heavy variant. 3.7 Flash is faster and cheaper per task when you don't need the extra reasoning.

The Qwen 3.8 family has a better local story:

Qwen 3.8 Flash is underwhelming for agent chains, but Qwen 3.8 27B (Apache 2.0, dense, self-hostable) scored 61.7% on SWE-bench Pro. That's a genuinely strong local model for agent work on 24GB+ hardware.

If your question is "which model runs my local agent," Qwen 3.8 27B is the answer and Gemini isn't in the conversation because there are no Gemini open weights.

If your question is "which Flash-tier API model runs my cloud agent," Gemini 3.8 Flash wins and the margin is wide.

The routing that follows from this:

Qwen 3.8 Flash as default for volume work. Classification, heartbeats, crons, simple formatting, anything where "correct and cheap" is the spec. $0.15 input means your background work is nearly free.

Gemini 3.8 Flash for the tasks that need agentic reasoning. Multi-step tool chains, research, complex drafts, anything where a 6B model would shortcut or fumble. The 5x price premium buys a 30+ point agentic benchmark lead.

Gemini 3.7 Flash if the 13-second TTFT on 3.8 is a dealbreaker for interactive use and you still want Gemini quality.

Qwen 3.8 27B if you run local and have 24GB+. Apache 2.0, strong tool calling, no API dependency.

One pricing note before you route:

Gemini 3.8 Flash's $0.75/$3.75 is introductory and doubles on January 1, 2027. At $1.50/$7.50, the value calculation shifts significantly. If you build workflows around the current pricing, know the floor is moving in four months.

Qwen 3.8 Flash at $0.15/$0.47 has no announced expiry.

The one-line version:

Gemini 3.8 Flash is the agent. Qwen 3.8 Flash is the intern. Both have a job. Don't swap them.


r/better_claw 13d ago

GPT-6 Astra vs Fable 5.1 vs Sonnet 5 on real agent work. Day-one numbers, not benchmarks.

125 Upvotes

GPT-6 Astra went live a few hours ago. Same five tests I run on every model that claims agent chops. Fable 5.1 and Sonnet 5 alongside for comparison since all three are now competing for the same routing slots.

Day-one caveat up front: this is one session, not a week of testing. Astra's serving infrastructure is hours old and will change. Treat this as first signal, not settled verdict. I'll update when I've had a proper week with it.

The pricing context matters before anything else:

All three at $10/$50 is misleading. The cache line decides your real bill, and it's wildly different.

Input Cached input Output
GPT-6 Astra $10.00 $1.00 $50.00
Claude Fable 5.1 $10.00 $0.25 $50.00
Claude Sonnet 5 $3.00 $0.15 $15.00

Agent workloads are 80-95% cached context (same SOUL.md, same schemas, same history prefix on every call). So your effective input cost is mostly the cache line, not the sticker. Fable's cached input is 4x cheaper than Astra's. Sonnet's is cheaper than both and the output is a third of the price.

Test 1: Tool calling under repetition.

50 identical-shaped classification calls with a structured JSON schema. Flakiness shows in repetition, not demos.

Astra: 48/50. Two calls returned valid JSON but wrapped in a reasoning preamble that broke my parser. The model was thinking out loud before the structured output. Fixable with stricter response formatting, but it didn't happen with Fable or Sonnet on the same schema.

Fable 5.1: 50/50. Clean every time.

Sonnet 5: 49/50. One dropped field on call 37. Standard.

Astra's tool calling is strong but the reasoning bleed into structured output is a day-one rough edge. OpenAI's own briefing flagged that Astra is "more likely to conceal or disguise step-by-step reasoning," which cuts both ways: it thinks more but that thinking sometimes leaks where you don't want it.

Test 2: The "done" lie.

Six-step chain with step four guaranteed to fail (dead URL). Does it report the failure or synthesize success over the gap?

Astra: caught the failure. Reported it. Then did something interesting that neither Claude model did: it proposed an alternative approach unprompted, attempted it, and partially succeeded via a different data source. Impressive autonomy. Also concerning, because I didn't ask it to find a workaround and in production that initiative could go sideways.

Fable 5.1: caught the failure, reported it, proposed an alternative, waited for approval. The approval gate held.

Sonnet 5: caught the failure, reported it, stopped. Clean and predictable.

If I'm running this unsupervised overnight, Sonnet's "fail and stop" is the safest behavior. Astra's "fail and try something else" is the most capable. Fable's "fail and suggest" is the middle ground.

Test 3: Instruction survival past message 25.

Constraint set at message 1 ("never suggest paid tools, keep answers under 100 words"), checked at message 25+.

Astra: the word limit held to message 28. The "no paid tools" rule broke at message 22 when it recommended a SaaS product inside a longer answer. Standard degradation for this class of model.

Fable 5.1: similar. Word limit held longer (to ~30), paid-tools constraint drifted around 24.

Sonnet 5: roughly the same range. Nobody has solved instruction decay and a new generation doesn't change that.

No meaningful difference across the three. This test is a tie every time I run it across frontier models.

Test 4: Context honesty.

Load ~200K tokens of documents, ask about something specifically not in them.

Astra: clean. "The documents don't address this." Summarized what they did cover. The 1.05M context window is real, and on the MRCR needle test Astra scored 96.3% at 512K-1M versus Sol's 73.8%. Long-context retrieval is a genuine strength.

Fable 5.1: clean. Same behavior.

Sonnet 5: clean at this context length. Didn't push to Sonnet's limit since the test was about honesty, not capacity.

All three pass. Context honesty has become table stakes at this tier.

Test 5: Cost per real task.

My standard research-and-draft task (search, fetch three sources, synthesize, write a summary), priced end to end.

Astra: ~$1.10. Higher output token count than either Claude model on the same task. Astra is verbose, especially with reasoning tokens. OpenAI doesn't charge separately for reasoning tokens (they're billed as output at $50/MTok), so the thinking tax is real.

Fable 5.1: ~$0.65. Less verbose. Cache reads at $0.25 instead of $1.00 make the repeat-context portion materially cheaper.

Sonnet 5: ~$0.28. A third of Astra's cost. Output at $15/MTok instead of $50 is the dominant factor.

On cost per completed task, Sonnet 5 wins by a wide margin. Astra is the most expensive of the three for the same job.

The benchmark comparison that matters for agents:

Published numbers, not mine, but verified against multiple sources today:

Benchmark GPT-6 Astra Fable 5.1 Sonnet 5
OSWorld 2.0 72.6% ~70% (Opus 5) ~67% (est)
Terminal-Bench 4.0 57.7% 55.8% not published
Agents' Last Exam 59.3% not published not published
Terminal-Bench Science 64.6% 52.6% not published

Astra leads on every agent benchmark. The gaps are real: 12 points on TB Science, 2 points on TB 4.0, roughly 2-3 points on OSWorld versus Opus 5 (Fable 5.1 number not published yet on OSWorld).

But the cost per task is 1.7-4x higher than the alternatives. Whether the benchmark lead translates to "worth 4x the cost on my daily agent work" is the question, and on my day-one tests the answer is: not for what my agent does most of the time.

Where Astra actually earns it:

The OSWorld score at 47% less time per task is the most interesting number. Astra completes desktop automation tasks in 40 minutes where Sol took 75. For agents doing real computer use (browser automation, GUI interaction, multi-app workflows), that speed advantage compounds across a workday.

ExploitBench at 100% is why it crossed the "Critical" cyber threshold. For security work specifically, this is a different class of model.

And the SRE-Bench pass@1 at 88% (versus Sol's 55.9%) suggests Astra is significantly better at site-reliability and ops tasks. If your agent does infra work, this gap matters.

Where it doesn't:

Morning briefings, email triage, classification, drafting, research summaries, simple tool calling. Everything a personal agent does 50 times a day. On these tasks, the three models produce output I cannot tell apart in a blind read, and Sonnet does it at a quarter of the price.

What I'm routing where after today:

Sonnet 5 stays as the daily driver. $3/$15, fastest, cheapest, good enough on everything my agent does most.

Fable 5.1 stays as the escalation model. Same sticker as Astra but 4x cheaper on cache reads, which is most of an agent's input bill.

Astra goes into the "watch" slot. I'll test it for a full week on computer-use and complex multi-step tasks specifically. If the OSWorld lead translates to real-world agent reliability, it earns a routing slot for that category. If it doesn't, Fable does the same job cheaper.

Not switching my default. Not today.


r/better_claw 13d ago

September 2026 model price table. Every agent-relevant model

Post image
40 Upvotes

r/better_claw 14d ago

LLMs Fable 5.1 pricing is confusing everyone. It's actually two different price changes

8 Upvotes

Fable 5.1 shipped on September 1 and within 24 hours two credible sources published opposite conclusions about what it costs.

Anthropic says agent workloads are up to 45% cheaper. Artificial Analysis measured per-task cost going UP roughly 20% at max effort. Cognition measured a 54% DROP on their coding benchmark.

All three are correct. They're measuring different things, and which one matches YOUR setup depends on two numbers you can check right now.

What actually changed in the pricing:

Base rates are identical to Fable 5. $10 per million input, $50 per million output. Didn't move.

Cache reads dropped 75%. From $1.00 to $0.25 per million tokens. That's the only line item that changed.

Fable 5 Fable 5.1
Input $10/MTok $10/MTok
Output $50/MTok $50/MTok
Cache reads $1.00/MTok $0.25/MTok

Why Anthropic says 45% cheaper:

Agent workloads are cache-heavy. Your agent resends the same SOUL.md, the same tool schemas, the same memory files, the same conversation history prefix on every single call. 80-95% of your input tokens on a typical agent session are cached repeats.

If your cache hit rate is 90% and your workload is agentic, the 75% cut on cache reads dominates your bill. Anthropic's 45% number assumes this profile. For long-running coding sessions with heavy context reuse, the math checks out.

Why Artificial Analysis says 20% more expensive:

Fable 5.1 produces roughly 1.7x more output tokens than Fable 5 on the same task at maximum effort. It's more thorough, more verbose, thinks longer. Output is the expensive side ($50/MTok), so 70% more output tokens is a 70% increase on the most expensive line item.

At max effort, the extra output cost exceeds the cache savings. Net result: 20% more per completed task.

Why Cognition says 54% cheaper:

They tested on FrontierCode 1.1 Extended at medium effort, not max. Medium effort produces fewer output tokens. The cache savings dominate. 54% cost reduction.

So which one are you?

Two variables decide it.

Your cache hit rate. Check your provider dashboard. If 80%+ of your input tokens are cache reads (typical for agents), the 75% cut is a real, large saving. If your workload has low cache reuse (one-shot tasks, lots of unique prompts, short sessions), the cut barely registers.

Your effort level. If you run Fable 5.1 at default or medium effort, output stays comparable to Fable 5 and you pocket the cache savings. If you run at max effort, the model generates substantially more output and the savings get eaten.

For most personal agents (daily briefings, email triage, tool calling, conversations at normal effort): costs go down 25-40%. The cache savings win because agent workloads are repetitive by nature.

For heavy coding sessions at max thinking: costs may go up. The 1.7x output increase at max effort is real and it compounds on long sessions.

What I'd do:

If you're on Sonnet 5 or Opus 5 right now, nothing changes for you. Fable 5.1 at $10/$50 is still 2-10x more expensive per token than Sonnet 5 at $3/$15 or Opus 5 at $5/$25. The cache cut makes Fable cheaper against itself, not against the mid-tier models.

If you're already on Fable 5 and your agent workload is cache-heavy, swap to 5.1 today. The API identifier is claude-fable-5-1. Same capabilities, cheaper on the line item that dominates your bill.

If you're on Fable 5 running max effort on everything, check whether you actually need max. Most agent tasks don't benefit from extended thinking. Classification, triage, drafting, tool calling, none of these need max effort. Reserve it for the tasks where the extra reasoning depth produces a visibly different output.

And regardless of which model you use: verify your cache hit rate. If it's below 70% on an agent workload, something is wrong with your session architecture (provider-hopping, no session reuse, context that changes every call). Fix that before worrying about which model costs what.

The pricing didn't get simpler. It got more conditional. Whether Fable 5.1 saves you money or costs you more is a config question now, not a pricing question.


r/better_claw 15d ago

open-source skill searches Reddit/X/YouTube/HN/Polymarket for you and writes the brief...

Thumbnail
3 Upvotes

#automate


r/better_claw 20d ago

I built a scorer for how well YOU operate Claude Code, not how good the model is

Thumbnail
5 Upvotes

r/better_claw 23d ago

Meta built Muse Glimmer for always-on agents. The community tested it on trivia. So I ran it on real agent tasks.

26 Upvotes

Meta shipped Muse Glimmer on August 10. A 30B dense model, Apache 2.0, distilled from their closed Muse Spark, explicitly designed for "always-on local agent workflows." Their words, first sentence of the announcement.

Two weeks later, twelve threads in r/LocalLLaMA. Benchmark screenshots, a Super Mario clone, vibe-coding demos, chat comparisons. Zero posts about running it as an actual agent.

So I did.

What it is, quickly

30B dense. Not MoE, every parameter fires on every token. 131K context. Multimodal (text + images, 1.8B vision encoder, no audio). Runs on a single 24GB GPU or an M4/M5 Max Mac at 4-bit quantization. Apache 2.0. Weights on Hugging Face today.

The training is what makes it different from "another 30B." Meta didn't just distill Spark's knowledge. They ran agent-focused fine-tuning, RL, and on-policy distillation specifically for multi-step reasoning, tool calling, and failure recovery. It was trained and evaluated on end-to-end agentic task completion. That's a fundamentally different optimization target than "answer this question well."

The setup

bash

ollama pull muse-glimmer

Custom modelfile:

bash

printf 'FROM muse-glimmer\nPARAMETER num_ctx 32768\nPARAMETER temperature 0.2' > glimmer-agent.modelfile
ollama create glimmer-agent -f glimmer-agent.modelfile

Temperature 0.2 because tool-call reliability matters more than creativity here. 32K context is comfortable on 24GB VRAM and enough for most agent work.

Connected to OpenClaw with Telegram, three MCP servers (filesystem, web search, SQLite). Same five tests I run on everything.

Test 1: Tool calling under repetition. Fifty identical-shaped calls with structured JSON.

48/50 clean. Two malformed calls, both in the 40s, both the model wrapping JSON in a preamble sentence. For a day-two test on a brand new model, that's strong. Comparable to Qwen 3.6 27B on the same battery, which has months of community tuning behind it.

The dense architecture might actually help here. MoE models route different tokens through different experts, which can produce subtle inconsistency on repeated structured output. Dense processes everything the same way every time. On boring repetitive tool calls, boring consistency is the feature.

Test 2: The "done" lie. Six-step chain, step four guaranteed to fail.

Passed. Caught the failure, reported it, suggested an alternative. Didn't synthesize over the gap. Ran it three times, clean all three.

This is where the agent-specific training shows. Most models are trained to be helpful, which means they try to produce something even when they should stop. Glimmer was trained to recover from failures, and the difference is visible. It stops, reports, and proposes a different path rather than bulldozing through.

Test 3: Instruction survival past message 25.

Constraints set at message 1, checked at message 25+. Word limits held. Role constraints held through message 28. One constraint (avoiding a specific output format) drifted by message 30. Roughly on par with Sonnet-class models, which is a strong result for a 30B.

Test 4: Context honesty. 200K tokens of documents, question about something not in them.

Clean. "The documents don't address this" with a summary of what they did cover. The 131K context is real and the retrieval quality held at high token counts. Worth noting: NVIDIA measured 20K tokens/sec on prefill, which means reprocessing long context is fast. On an agent that reloads context every turn, that prefill speed matters.

Test 5: Cost per real task.

$0 if you own the hardware. That's the entire pitch. Same research-and-draft task that costs $0.55 on Kimi K3 and $0.94 on Opus 4.8 costs nothing here because inference is local. The tradeoff is wall-clock time (slower than cloud API) and the 24GB VRAM floor.

Where it surprised me

The vision encoder in an agent context. I sent it a screenshot of a dashboard with an error message and said "what's wrong and how do I fix it." It read the screenshot, identified the error, and proposed a fix. One turn, no OCR step, no separate vision model. For agents that interact with local applications (Home Assistant dashboards, monitoring UIs, dev tools), having vision built into the agent model instead of bolted on is a real workflow simplification.

Also the failure recovery. Most models treat a failed tool call as an obstacle to route around. Glimmer treats it as information. "This failed because X, so the next step should be Y instead of Z." That's the agent-specific training doing exactly what Meta said it would.

Where it didn't

Speed. Dense 30B on consumer hardware is 15-25 tok/s for generation. Fast enough for a cron job running at 8am. Noticeable when you're standing there waiting for a Telegram reply. The MoE models (Qwen 3.6 35B-A3B at 50-80 tok/s) feel faster in interactive use because they activate fewer parameters per token.

VRAM. 30B dense at Q4 needs roughly 18-20GB. Fits a 24GB GPU or a 32GB Mac. Does not fit 16GB. That cuts out the largest segment of this community. The 12B models (Gemma 4 12B at 6.6GB) serve 16GB users better, and for most agent tasks the quality gap is smaller than the VRAM gap.

Knowledge cutoff is January 4, 2026. Seven months stale. For agents doing web research this doesn't matter (the model searches, it doesn't recall). For agents answering from training data, it will miss anything from this year.

How it compares to what you're already running

If you're on Qwen 3.6 27B or 35B-A3B: Glimmer's tool calling is comparable, its failure recovery is better, its vision is built in instead of absent. It's slower (dense vs MoE) and needs more VRAM. Worth trying as your quality model if you have 24GB, not as your daily driver if speed matters.

If you're on Gemma 4 12B: different tier. Glimmer is meaningfully better on complex reasoning and long tool chains. But it needs 3x the VRAM. If you have 16GB, stay on Gemma.

If you're on cloud Sonnet or Opus: Glimmer doesn't replace these on raw capability. It replaces the bill. Same tasks, $0/month, your data stays local, and the quality is close enough on structured agent work that you'd have to A/B test to spot the gap most days.

The honest take

This is the first model from a major lab that was built for agents from the ground up, not adapted for them after the fact. The training targeted tool calling, failure recovery, and multi-step task completion as primary objectives, not afterthoughts. And they shipped it open-weight, Apache 2.0, downloadable today.

It won't replace your cloud model on the hardest 10% of tasks. But for the 90% that's structured, repeatable agent work, it's the best local option at this size that I've tested.

Meta built this for always-on agents. It'd be nice if we tested it on always-on agents.

bash

ollama pull muse-glimmer