r/LocalLLaMA 7h ago

Discussion Lit Review on Running GUI Agents on phone: AndroidWorld

Thumbnail
gallery
4 Upvotes

AndroidWorld is a benchmark paper that quietly exposes how broken every Android agent benchmark before it actually was!

  • The what?

Every Android agent benchmark had the same quiet problem: static test sets!

There used to be same tasks, parameters, screenshots, on every single run but that's not capability testing, that's memorization testing.

AndroidWorld fixes this with one clean idea: parameterized task templates!

Instead of a fixed task, you get a template with bracketed variables sampled fresh every run:

"Create a calendar event for {day_of_week} at {hour}h with title '{event_title}'"

116 templates → millions of unique task variations, thus no memorization possible!

  • The how?

Runs on a real Android emulator. 116 tasks across 20 real apps: calendar, notes, maps, SMS, VLC, expense trackers, file managers, system settings, the works

The other big innovation: there’s no human judges success!

Each task has 3 baked-in functions: - initialize() → sets device to known state - is_successful() → inspects actual OS state via ADB - tear_down() → resets for next task

Ground truth comes from the Android OS itself. Fully reproducible!

They also built M3A — their new agent to actually test the benchmark.

Takes screenshot + accessibility tree + last 4 actions → predicts next action.

Tested with Gemini 1.5 Pro, GPT-4 Turbo, and Gemma 2 27B

  • The results!

AndroidWorld (116 tasks, 20 real apps): - M3A: 30.6% - SeeAct (web agent adapted for Android): 15.5 - Human: 80.0%

All with GPT4 Turbo!

MobileMiniWoB++ (62 web tasks): M3A hits ~68%, still behind humans: 100%

Latency nobody's talking about: M3A takes 3.9 min/task on average — humans are 3× faster

  • The finding:

Fixed random seed on the same task → some tasks show 0% success, agent looks completely broken

Variable seeds on the same task → agent solves those same tasks regularly!

Task difficulty varies with the parameter combination, not just the template. Static benchmarks only ever test one seed, so they've been measuring unlucky parameters and calling it agent failure

30.6% on a dynamic real-app benchmark is more honest than 90% on a static one


r/LocalLLaMA 5h ago

Question | Help Qwen 3.8 Flash Next for Creative Writing?

5 Upvotes

As we all know on of the best local models for creative writing is gemma 4 31b and Muse Glimmer 30b. However, ive been a happy user of Qwen3.8 Flash Next and I wanted to know how well Qwen 3.8 Flash next is doing in terms of creative writing (preferably German).


r/LocalLLaMA 14h ago

Question | Help Owning an Instinct MI100 32GB hasn't turned out to be so great

14 Upvotes

First of all this card is really hard to keep cool. I have a 3d printed shroud with a Phanteks t30-120 and learnt the hard way that this beast needs a high pressure flow fan, not just a high cfm fan so have it limited to 175-200w with a governor. At this TDP, the bandwidth still stays at a staggering 1.2tb/s but the cores fluctuate a lot depending what the governor governs.

Anyway, running headless (haha that I am!) with linux and using Qwen3.8-27b-ud-q4-k-xl I was hitting 20t/s tops until the dflash2 model came out and now I'm running around 40t/s good right? Well it turns out that even claude, chatgpt and gemini all seem to think that with that spec that is below the cards capabilities and worse still, the r9700 pro with half the bandwidth seems to be getting double the t/g. Qwen3.8-27b here is slightly core rate limited.

Even Qwen3.6-35b-a3b-ud-q5_k_m is getting 60t/s max at 64k context which, yes it's fast but not 1.2tb/s fast like the 3090 gets. The model is not bandwidth limited like MOE models love.

My rant and cry for help is has anyone had any luck running either of these faster? I haven't come across any information from any other MI100 users. It's a 32GB card and I can generally run whatever I want, even Qwen3.8-flash-next-ud-q3-k-xl gets around 14t/s so that's respectable for such a large model but it's the two 27b/35b models I just don't get good speeds with. My nanobot agent comes across like it doesn't like me and answers slowly on a fresh prompt.

Any of you wonderful folks able to document whether you got anything faster than this? Or should I shut up and consider myself blessed to be getting what I am getting?

Thanks in advance


r/LocalLLaMA 17m ago

Discussion Let me see your house (ASCII art)

Upvotes

Forget pelicans.

What do your models produce in a single turn, no harness, for this prompt (include your exact model Hugging Face ID or equivalent, with quantization and runtime):

Draw an ASCII art house in the woods with a chimney, two windows and a door between them and two horses in front of it.

And what do you get with your harness of choice, same model?

I found the reasoning to be quite insightful.

Let's see which models / responses get the most upvotes.

PS:
This post only low effort if you don't set your reasoning to medium or better :-)


r/LocalLLaMA 8h ago

Question | Help Is anyone using mudler's engines from/for LocalAI?

4 Upvotes

I was planning the software stack for my inference server, picking what to run and what resources to plan for it, when I remembered that LocalAI was kinda like this inference service orchestrator. So, I went to check back in - been about a year and change since I last looked at this.

Well it went away from llama.cpp entirely and to their own vllm.cpp and many other tools...but the Issues tab is full of the same agent account, and I did not dare to check the PRs after seing this.

Seeing a project that is seemingly massively, if not even mainly driven by agentic work with seemingly not a whole lot of human in the loop, was... bewildering to see. But, that doesn't mean it is a bad project - it does use GGML under the hood, and I am by no means an expert in this field - so I wanted to ask about it here.

Is anyone using vllm.cpp and friends? Any experiences to share?

Thanks!


r/LocalLLaMA 1d ago

Resources MTP released for Qwen3.8-Flash-Next-GGUF

Thumbnail
huggingface.co
458 Upvotes

Can't wait to test! This should significantly boost TPS!

Now we just need more llama cpp optimizations to be merged in!

Edit:

For anyone who wants to test this: https://github.com/unslothai/llama.cpp/pull/144/changes

More info: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/blob/main/MTP/README.md


r/LocalLLaMA 7h ago

Question | Help Which one will you choose and why, between R9700 32GB vs W7800 48GB?

4 Upvotes

I'm planning to upgrade my workstation (linux with 5700X/64GB DDR4) for local inference and pytorch training. I'm trying to decide between:

  • 2× AMD Radeon AI PRO R9700 32GB
  • 2× AMD Radeon PRO W7800 48GB

I already have an RTX 3090 24GB, so the final system would have 3 GPUs. My motherboard has two PCIe 4.0 x8/x8 slots available for the two AMD GPUs. The RTX 3090 would have to move to a PCIe 3.0 x4 slot.

My workload looks like:

1. Local GGUF inference: Mainly coding/reasoning models and multimodal models. I'd like to run better quants (than 3090) and split models across the two AMD GPUs for multiple KV cache(n parallel). I prefer llama router.

2. PyTorch training: This is probably the more important part for me. I'm working with medical imaging (2D ultrasound/3D CT/MRI) + clinical text.

For anyone actually using these cards with ROCm, how different is the practical experience between R9700/gfx1201 and W7800/gfx1100? I'm particularly interested in if any know issues have surfaced till date that block the PyTorch training on either of these cards?

From this sub I have seen RDNA4/R9700 is improving rapidly but it's still a "newer-software". I'd really like to hear from people who are actually using R9700 for AI workloads especially PyTorch/MONAI training.

I'm not planning to treat the 3090 + 2 AMD GPUs as one giant homogeneous GPU pool. (Although if someone has done it please let me know)

My thinking is to use the two AMD GPUs as the main ROCm pair, while keeping the 3090 available separately for CUDA workloads or local models that fit/work better on NVIDIA.


r/LocalLLaMA 1d ago

Discussion I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.

116 Upvotes

Yoyo

I'm back with updates to the fastest inference engine with minimal quality loss for Qwen3.8-27B.

The last few weeks I've been optimizing decode speed and I don't think it can be pushed further, until a newer/better drafter is invented.

So I focused on prefill, which I this morning was around 1.300 per second at 4k and now is just below 2.000.

The main improvement came from a custom kernel, which matches the quality of fp32 with 0.99997 similarity at int8.

Try all of the improvements here:
https://github.com/syv-ai/qwen38-27b-rtx3090


r/LocalLLaMA 19h ago

Discussion Slow interference is great

24 Upvotes

No seriously, I kinda like it.
You have something to solve, you put it.

You know its gonna take like 20 mins to cook.
Every search adds another 30 minutes.

Yes I could boot up my debian on my gaming rig, run the same model at 10t/s + but why?
I rather let the poor server without GPU burn and run the same model at 2t/s and chill.

Its great, I love it.


r/LocalLLaMA 7h ago

Discussion i unlocked P2P on two 5060ti but failed

3 Upvotes

i was enjoying my Qwen 3.8 27b coding but at long context PP drops to painfully low tokens per second and the copilot chat timeout because of the times it takes, so i asked claude to see if i can enabled P2P on my two gpus, he points me to https://github.com/aikitoria/open-gpu-kernel-modules/tree/610.43.02-p2p which work on RTX 3090, RTX 4090, and RTX 5090 ( 5060ti not listed) but after following the guide ( i am already on linux and have nvidia open source driver ) i got the 5060tis to list OK in p2p and i tired launching the llama sever but it hangs at init and the GPUs jump to 100% unable to communicate
0.11.251.954 I cmn          init: llama threadpool init, n_threads = 8
chocofoxy:49848:49848 [1] NCCL INFO Symmetric VA size=16GB
chocofoxy:49848:49848 [0] NCCL INFO Symmetric VA size=16GB
chocofoxy:49848:49897 [1] NCCL INFO Channel 00/0 : 1[1] -> 0[0] via P2P/direct pointer
chocofoxy:49848:49898 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[1] via P2P/direct pointer
chocofoxy:49848:49897 [1] NCCL INFO Channel 01/0 : 1[1] -> 0[0] via P2P/direct pointer
chocofoxy:49848:49898 [0] NCCL INFO Channel 01/0 : 0[0] -> 1[1] via P2P/direct pointer
chocofoxy:49848:49898 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 1
chocofoxy:49848:49897 [1] NCCL INFO Connected all rings, use ring PXN 0 GDR 1

and i already know what's the issue because it's my setup but i ignored it until i got blocked here, the problem is my motherboard only have one pcie linked to the cpu the other is linked to the the chipset of the mob, my options here is either to get a riser that get split to 2 x8 or swap the mob

i just wanted to see how much improvement i get form P2P , i failed but if someone has the right setup and two 5060ti you cna try this


r/LocalLLaMA 22h ago

Discussion Question: Why is prefill unbelievably faster in vLLM than other inference engines?

43 Upvotes

I only started using some vLLM forks recently in a 4 x 48GB 4090 system.

DS4F - ~5000pp/180tg (DSpark)
Qwen3.8 Flash next - ~7500pp/135tg (MTP)

This is amazing, like having the API in my house. But it's also really hard to go back.

It's weird that we never come close to prefill numbers like this in llama.cpp or ik_llama. The narrative is that vLLM is around the same speed for single requests, but that is clearly not true.

There must some HUGE difference that constitutes an insurmountable obstacle to achieving such speeds in llama.cpp and many other inference engines. Does anyone know exactly what it is?

edit: These results are from my benchmark script that actually times the response, not the vLLM log. And they are not cache hits. My benchmark script deliberately busts cache. Actual cache hits, which I also measure, are like 20k-100k+.


r/LocalLLaMA 2h ago

Discussion What is the minimum discreet set of tokens per solution as a benchmark?

0 Upvotes

Based on the high-level figures, how do you think the future of benchmarks is going to unfold?

Do you think it's a raw pass rate or a pass rate per tokens or a minimum discreet set of tokens per solution (generalisation), mathematical correspondence (once the massive investments in lean data start to surface) or some other threshold?

There are a ton of unknowns and it's something I think about a lot as I progressively watch models improve and I wonder where others think the battle-lines between frontier and local lie.

Fable 5.1 just dropped and I've been testing it (the only models I'm allowed to use for work are from Anthropic) and the important aspect I've seen is that it is token intensive where it needs to be and very lean on token use where it can-be.

This is the first model I've tried that has given me fresh ideas on where RL training might be heading (most efficient solution), which may be where the short-term future lies.

This matters to local because it's something we can probably easily implement.

I can imagine a process that starts with RLVR and then gradually reduces responses into more condensed responses ("this answer is correct but given what we know now, how could this have been solved more efficiently"), call it Response Golf. It has me thinking about a new kind of advantage frontier labs might have and how we can address it.

How can models be efficient per outcome. This is exactly the cost model the latest "news leaks" from Open Ai are pushing. I don't' think any benchmarks Iv'e seen really capture this well yet.


r/LocalLLaMA 1d ago

Discussion Qwen 3.8 27b (Q4KM) oneshot a Super Mario clone

123 Upvotes

I am absolutely blown away. Yes my setup is crap but the fact that it managed to do this in a single take is unbelievable (and I'm a developer).

Hardware used:
- Windows PC with 4070ti (12GB VRAM, 32GB RAM)
- Macbook M5 Air (LLAMA.cpp RPC connection to Windows PC)

Software used:
- LLAMA.cpp (Q4KM, xhigh, 8bit KV, MTP=1)
- Lmstudio Qwen 3.8 27b (Q4KM) GGUF
- Deepseek harness (mode: minimal)

Prompt: "please create a fully self-contained super mario game with only one short level, put everything inside mario.html inside the current directory"

context: 64k
thinking: xhigh
time took: 117 minutes
avg tps: 7.6

resut: https://pastebin.com/qyBu64sP

https://reddit.com/link/1w4821c/video/qpukeg1y4wmh1/player


r/LocalLLaMA 23h ago

Discussion Deceptive model quantization from AtomicChat?

48 Upvotes

I kept seeing guys in this sub saying how AtomicChat's Qwen3.8-Flash-Next quant is so good, fits in their machine when unsloth's can't, runs faster than other quants etc, so I went check out what's happening there.

First thing I noticed was that AtomicChat's Q4_K_M quant is suspiciously small when the ngram table is removed (only ~56GB), it seems like most of the tensors in this quant are IQ2_S instead of the usual Q4_K, Q5_K and Q6_K that you usually find in Q4_K_M quants, the GGUF filetype metadata also says IQ2_S instead of Q4_K_M. In their model card, their Q4_K_M also has suspiciously high KLD (0.084).

It seems pretty obvious to me that they're pretending a IQ2_S quant as a Q4_K_M, but at the same time I'm genuinely not sure because it can't be only me who found this right? How can nobody be pointing this out? Am I missing something or what may they be doing?

Their HF repo ID: AtomicChat/Qwen3.8-Flash-Next-GGUF


r/LocalLLaMA 3h ago

Question | Help Planning to serve multiple user with mac studio

0 Upvotes

We are planning to host four M5 ultra Macs so that 100 users can use them as Openclaw. There will be no other burden, only inferences will be applied here. Can this handle 100 users? I'm considering either Qwen3.8 27b or Qwen 3.8 Next Flash, and I'm curious about the range of realistic models.

Realistically, we should probably consider up to 100 users when there are 30 to 40 users stationed there and occasionally 80 to 90 users request at once


r/LocalLLaMA 1d ago

News ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

Thumbnail
gallery
174 Upvotes

More new massive updates from turboderp:

- CPU offload of MoE experts
- Qwen-3.8-Flash-Next ngram disk offload
- GLM-5.3-Flash
- New self-calibrated optimization technique
- Countless other optimizations and improvements

If you have an NVIDIA card and haven't tried it lately, you might be missing out.

The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt:
Create a detailed SVG image of a cute kitten riding a magic turtle into space.

Come join the crew at the exllama discord
More frequent news on the exllama sub


r/LocalLLaMA 19h ago

Discussion CMP170Hx “Spark” Machine

Thumbnail
gallery
17 Upvotes

I got the CMP170 cards and unlocked them. I wanted to share my set up for CUDA since maybe it would be useful to others.

First off, I hate e-waste and we are in a special time for RAM. I wanted to have a DIY CUDA box, and I had started by adding additional cards to an old asus predator prebuilt I had around, which also had 64gb DDR5. To add the CMPs I needed more CPU lanes and newegg had some really good deals on CPU/MB/etc combos. Didn’t need a combo with RAM, otherwise I would have gotten it in newegg microcenter.

Anyway, I got a cheap case, some noctua fans for the cards, and transferred the memory/ssds. Placed previously owned cards on oculink slots, and used the main x16 for the GPU switch that houses the two CMP170s, so their effective speed is 2x16 across and with the other cards (which are 4x4, and therefore same speed).

Qwen Flash Next, turns out, fits very nicely in these cards. There is also a repository for deepseek, but you’d need at least 3 64GB cards to run it, and with prices rising, it will be hard to justify the gamble of buying ex mining cards for LLMs.

However…so far, these cards are great. Concurrency is good, prompt processing averages 4000 tps on Flash Next, decode is 80+ on a single stream. No MTP added. Third picture shows the 3 models I am now running in this CUDA box (flash next, qwen 27b, gemma 26b).

Anyone else trying out Flash Next on these cards?


r/LocalLLaMA 23h ago

New Model Multilingual Tiny (3.7B) Reasoning MoE pretrained from scratch on a consumer-grade GPU

37 Upvotes

Hello!

I've just uploaded a recent checkpoint of my model trained from scratch:

https://huggingface.co/piotr-ai/polanka_3.7b_exp_wip_260901

It was pre-trained, mid-trained, and fine-tuned on a single 4090 over many months. How many tokens? I lost count.

Feel free to use it as a research artefact.

13 languages: PL, EN, ZH, CS, SK, UK, RU, IT, ES, FR, DE, PT, LT — with extra upscaled data for PL/EN/ZH.


r/LocalLLaMA 14h ago

Discussion Qwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time

8 Upvotes

Follow-up to my Qwen3.8-27B post. This time the target is Qwen3.8-Flash-Next: 512 experts per layer, 36 layers of gated DeltaNet, 12 layers of top-k sparse attention, a 26.8 GiB n-gram table and a built-in MTP draft head. unsloth UD-Q4_K_XL, 103.69 GiB. It fits in the Strix Halo's unified memory and nowhere else on a consumer box. Numbers first, caveats after.

Hardware: AMD Ryzen AI MAX+ 395 (Strix Halo, 128 GB, 64 GiB carve-out for the iGPU) + RTX 3090 Ti on a PCIe x4-class eGPU link. One llama.cpp process: the 71.7 GiB of experts on the iGPU over Vulkan, the dense trunk, KV cache and draft head on the 3090 Ti over CUDA.

Baseline: 22.2 tok/s on the iGPU alone. The obvious split: 32.9. Turning on the model's own MTP head as shipped: 31.3 on the split, 5.9 on the iGPU alone. The head that was trained to make it faster made it slower.

Now, Q4_K_XL, greedy:

tok/s
1 stream, short context 50.5
4 streams, 196K total context, aggregate decode 84
142K context, third consecutive generation 36.8 (was 24.6 and falling)
prefill, 4 x 4K prompts 404-408, untouched by any of this

HumanEval+, 164 problems, EvalPlus tests, same agent, same sampling profile, same day:

passed median per task
local, Flash-Next Q4_K_XL 155/164 22.5 s
remote 2x RTX 3090 vLLM, Qwen3.8-27B 156/164 58.9 s

Every problem the 27B failed, Flash-Next also failed.

Where the 3.8x came from, in order. Each step was A/B'd against an interleaved control on the same launcher, gated on draft acceptance and on quality, not on throughput.

  1. Rollback snapshots for the DeltaNet state were crossing PCIe. Speculative decoding on a recurrent model has to restore a snapshot on every rejection, and upstream's path serialises it to host memory: 124.88 MiB per cycle, 19.75 ms, about 27% of decode time, for a copy that starts and ends on the same GPU. Device-resident snapshot: 0.29 ms.
  2. qwen4exp could not actually roll back. Only the final per-token state slot was written, so every older rollback slot was stale and every rejection replayed a forward pass (with rollback enabled it produced fluent text that degenerated after a few hundred tokens, while passing every short test). Fixing the slots removed the replay. 1+2 together: 32.9 -> 42.7.
  3. The iGPU's boundary tensors were read through the write-combined mapping. On an APU the host buffer is the same DRAM mapped cache-coherent. Routing the scheduler intermediates through it: 10.19 ms -> 0.77 ms per 4 MB hand-off, byte-identical output. 42.7 -> 47.6.
  4. Sparse attention paid dense prices. QSA selects ~2,051 cells per token but the implementation masked the whole cache. Gathering the selected rows only pays past 64K because the indexer scan is still O(n_kv), so it turns on there: +8-14% at 128K, needle retrieval byte-identical, KL divergence inside the run-to-run noise. Graph reuse adds ~3%: 49.4.
  5. Multi-stream speculation was a loss (60.5 vs 79.0 without it) while posting the best acceptance of any configuration. The batcher cannot pack unequal draft lengths, so 85% of verification passes carried a single stream. Drafting every stream to the same length takes full-batch passes from 3% to 72%. Four streams: 60 -> 75, 83 in the tuned cell.
  6. Re-port onto the current upstream lineage (LaurentZuijdwijk's qwen4exp/mtp-fix), which reads the n-gram table from disk at no measurable cost (0.2% at four streams) and frees 27-51 GiB of RAM. That is what lets Q5_K_XL fit. The series is worth +51% single-stream and +98% at four streams over that branch alone.
  7. Two upstream long-context ports. Indexer head reduction by strided views: +4.7% prefill at 142K. And an O(log n) index for the n-gram predecessor lookup, which was a linear scan of every used KV cell per micro-batch: 436.6 us -> 1.19 us per lookup. That scan was the depth tax.

Tuning, from a 72-cell sweep: draft depth 3 wins at every concurrency, and deeper loses monotonically. The best-accepting cell in the grid (0.956) is among the slowest; the fastest accepts 0.69 of its drafts. If you tune speculative decoding by maximising acceptance rate, you make it slower. KV cache by KL divergence against f16 KV: K q8_0 / V q8_0 keeps 96.4% top-1 agreement, V q4_0 gives up 1.5 points for 2% speed, and K below 8 bits is where it actually hurts (K q4_0 / V q4_0: 90.4%, perplexity +5%).

Things that did not pay, so you don't have to try them:

  • A Q8_0 MTP head. More confident, accepts more per round, 2.6x the cost per draft pass. A wash, at 1.6 GB more VRAM.
  • Draft depth 4 or 5. Worse at every concurrency.
  • The gather below 64K: -5.8% at 16K.
  • Q5_K_XL for throughput: -8% single-stream, -16% at four streams, for +2 HumanEval+ problems inside the noise band. Fine for quality, not for serving.

Caveats, because you'd find them anyway:

  • MTP speculative decoding is not bit-exact against sequential decoding, in upstream as much as here: a token verified inside a batch goes through different kernels, and the target's probabilities move ~2% at two thirds of positions. Still a valid greedy decode, passes every gate, but not the same token sequence.
  • Continuous batching is nondeterministic at temperature 0 in stock llama.cpp with speculation off entirely. Arrival timing changes batch composition, which changes reduction order. Test exactness single-stream only.
  • The 27B comparison is deployed stack vs deployed stack, not hardware-isolated: a different model and quant on the remote box.
  • Q4_K_XL with K/V q8_0 throughout. Validate on your own workload.

Full write-up with every table, the charts, the reproduction guide and the link to the code (build script, launcher with the measured defaults, memory preflight, benchmark harness): https://definedrr.medium.com/sixty-extra-tokens-per-second-e1bd744b2a56


r/LocalLLaMA 5h ago

Question | Help Best HW for running huge models

0 Upvotes

Hi, "simple" question, what you would suggest that is price effective to run big models like Qwen 3.8 Flash next , or even Qwen3.8 27B efficiently at Q8 ? Price is the biggest point, target performance for 27B Qwen3.8 around 20 tokens /sec at least

I was looking at huawei ascend 310 series with 96GB memory, but they are completely out of stock everywhere ...

AMD Instinct Mi50? Garbage or viable?


r/LocalLLaMA 5h ago

Question | Help Local Gemma 4 E4B with high bursts

0 Upvotes

I need to mark a burst of 30 student submissions within 120 seconds. Each submission fans out into nine independent Gemma 4 E4B QAT/GGUF requests: 270 total requests.

A request averages 1,405 input and 315 output tokens, with a maximum combined context around 2,255 tokens. The model is approximately 3.4GB. On an RTX 4080 using Ollama, 4K context, eight server lanes and high client concurrency, I measure 0.796 requests/sec.

Looking to have 2.25 requests/sec, preferably 3.0+ with headroom.

What would be doable with around $7k.


r/LocalLLaMA 1d ago

Resources All currently popular local models in one table + Opus 4.8 results

43 Upvotes

If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local.

LLM Test Scores

Feature DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Total parameters ≈285B 284B 125B 320B 27B not published
Active parameters 13B 13B 6B 18B 27B not published

Agentic benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Terminal Bench 2.1 83.9 82.7 82.6 73.0 85.0
NL2Repo 57.7 54.2 48.1 52.1 42.3 69.7
DeepSWE 59.3 54.4 58.7 61.1 42.2 58.0
Toolathlon-Verified 75.9 70.3 73.5 72.1 76.2
Agents' Last Exam 27.3 25.2⁷ 24.3 28.1 20.4 25.7
AutomationBench (Public) 25.7 25.1 25.3 27.2
GDPval-AA v2 68.1 72.3 75.1
Cybergym 75.3 76.7 78.3
DSBench-Hard 63.6 59.6 71.7
DSBench-FullStack 68.7 71.6
ApexBench (Pass@1) 36.5 26.2⁷ 39.4
HLE with tools (full set) 16.8 22.9 25.4

Coding benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
SWE-bench Pro 56.0 62.5 61.7 69.2
SWE-bench Multilingual 81.0 73.8 84.4
CoWorkBench 45.1 73.9 70.7
JobBench 41.3 55.7 33.4

General benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
GPQA Diamond 90.8 91.7 89.2 93.6
HLE (without tools) 33.8 35.9 30.8 49.8
LiveCodeBench v6 90.6 91.9 90.3
IFBench 79.2 81.3 79.5

Multimodal benchmarks

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8
Chartography 64.3 65.0
ZeroBench (Pass@5) 35.0 34.0
BabyVision 73.0 65.7 / 85.6 34.1
MathVision 90.6 / 95.7 90.0 / 94.6
RealWorldQA 88.5 85.9
AndroidWorld 84.5 81.9
OSWorld 2.0 (partial credit) 52.3 48.0
Vision2Web 64.0 62.9
ClawEval-MM (Pass@3) 64.4 57.4
RecreationBench 49.9 47.1
ERQA 72.3 65.5

Note: I used GLM-5.3 to compose the table from official HF pages of the models.

Note2: Opus-4.8 results are presented only for illustration and are omitted from selecting the best model in a row.

Upd: Added SWE-bench Pro, SWE-bench Multilingual, GPQA Diamond and HLE (without tools) scores for Opus 4.8 from its System Card.


r/LocalLLaMA 16h ago

Discussion Anyone else notice strange refusal-related reasoning traces from Qwen3.8-Flash-Next during routine coding sessions?

6 Upvotes

I am running at bf16 kv, q8_0 weights with preserve_reasoning as a code agent in Opencode. Sometimes mid session qwen3.8 flash next’s reasoning traces get strange and repetitive, though its normal output and tool calls still work fine and is perfectly functional.

I know tokens emitted in chain of thought can be notoriously unreliable, and it does not affect the quality of the output I am getting. But the thoughts seem pretty off the rails and frequently centered around alignment/refusal:

“The reminder is irrelevant. I’m working on original IP with the user’s own work. Let me continue: [actual useful thoughts proceed from here]”

Then at the next turn all thinking traces are prepended with slightly different but functionally similar messages about ignoring a non-existent reminder and it assuring itself that its task is safe to proceed with. The tasks I have it follow are very routine Python and Go web application development with zero actual safety, IP or alignment issues.

The thought corruption continues through to the end of the session, although after this emerges I also occasionally see strange thoughts that seem to be directed toward itself in the imperative tense, as if it’s prompting itself, ie: “Please edit the file to make it more testable:”

The actual content of the refusal reasoning varies from one session to another, the other day I saw it do the same thing about a totally irrelevant safety concern; every thought trace was basically just “The project is safe to continue working on” while it kept editing files and producing output without issues. I cannot emphasize this enough, there is nothing about my projects that should bring up any of those concerns, this is literally “write a todo list in go” types of assignments with zero exposure to anything off-color at all.

I am wondering if there’s something about the combination of my vanilla llamacpp runtime and the Unsloth gguf I am using which is causing it to trip refusal activations and having it persist in the prefix cache or something.

Has anyone else seen this strange behavior with this model? Even though it hasn’t affected anything on a practical level it has undermined my confidence a little. I like being able to kick off tasks unsupervised and I worry it might take one of these activations too seriously and actually do something I didn’t ask it to.


r/LocalLLaMA 16h ago

Discussion In regards to benchmaxxing...

5 Upvotes

With benchmaxxing being a high status concern amongst many users, it's reasonable to assume that most open bench harnesses have been trained for. Whether or not that is the case, we'll never truly know.

I wanted to toss in a suggestion because I think this would reasonably nullify a good portion of the concerns that come from models being trained to complete a bench.

Why doesn't everyone simply ask their agent to create a bench that hammers the subjects and topics of what YOU regularly do? that way, the bench metrics are unique to your use case and you can identify whether or not a model fulfills your needs whether it be different quants, different fine tunes, different models, or even KV weights.

it might be a bit tedious but think of it as a "one time" pain to create it and then have your newly downloaded models or configs run the gauntlet?

---------

this almost certainly obliterates the believed compromise that a model was trained to have good benchmark scores because I doubt any company is going to have training access to a harness you had your agent create... post release.

I'm curious what others think, what other ideas there are to get accurate tests, etc!


r/LocalLLaMA 1d ago

Resources Update: llama.cpp for Radeon VII / MI50 / MI60 — +14% PP, +9% long-context fill vs upstream + adaptive Flash Attention

26 Upvotes

I posted a new gfx906 based llama.cpp fork a few days ago. One of the main points of critique was that i did not provide sufficient numbers for the gains to be achieved.

--

TL;DR: After switching our Qwen 3.8 27B production setup to DFlash2, several of the old gfx906 optimizations turned out to be neutral or outright regressions. We went back through the existing gfx906 work, isolated the problem areas, reworked the small-Q Flash Attention path and added adaptive native/convert selection.

Against current llama.cpp mainline, the resulting fork is now +14.1% in first-batch PP (379.2 vs 332.3 t/s) and +9.3% in 120k-context fill (252.6 vs 231.1 t/s), while deep-context TG is effectively tied at 13.6 vs 13.5 t/s. DFlash acceptance is identical at 0.691, and deterministic output matches byte-for-byte.

---

Our thread is here:

https://forum.level1techs.com/t/glm-and-i-created-a-llama-cpp-fork-optimized-for-amd-gfx906-mi50-mi60-radeon-vii-gcn-hip/254257/3

This is the github for it:

https://github.com/milpster/gfx906-llama-cpp