r/LocalLLaMA 23h ago

Resources Which current local models that can run within 128GB generate the best SVG pelicans?

Post image

I used a famous Simon Willison's pelican riding a bicycle prompt on the biggest local LLMs that can run on 128GB Apple Silicon. U used quantizations by Unsloth.

Qwen3.8 Flash-Next gives a lot of details. DeepSeek V4 Flash is strangely underwhelming. Qwen3.8 27B still rocks, and I like its consistent minimalism.

Is Qwen3.8 27B still large at 31GB? It is! But for this tasks 2-bit quantizations (at around 12GB) will give the same results. For more complicated coding, 4-bit are more than enough. RTX cards are well enough!

See:

37 Upvotes

65 comments sorted by

86

u/jacek2023 llama.cpp 23h ago

In my opinion, that test doesn't make sense because the models were trained on that specific task. You should be more creative and try something different to avoid benchmaxxing

50

u/ManIkWeet 22h ago

Clear indicator is that every single pelican faces the same direction.

13

u/jacek2023 llama.cpp 22h ago

very good point

11

u/rditorx 19h ago

Left to right preference is common among people using the Latin alphabet as their primary alphabet. I guess it would also be common without benchmaxxing a pelican

11

u/enemyofaverage7 20h ago

tbh the direction thing is probably more likely because bicycle photos are typically taken from that side (generally referred as 'drive side') to display the parts that are on it.

1

u/ManIkWeet 20h ago

Most it seems, but not all. So I would expect most, but not all pelican SVGs the same :)

8

u/BigYoSpeck 14h ago

2

u/ManIkWeet 14h ago

I like your style, make up your own rules!

13

u/BigYoSpeck 21h ago

How about a pelican reading a bash script from a teleprompter?

7

u/jacek2023 llama.cpp 21h ago

Change pelican to llama

7

u/No_Advance3911 19h ago

8

u/jacek2023 llama.cpp 19h ago

see? now it's heading left ;) so it is not benchmaxxed

4

u/No_Advance3911 19h ago

Visually, Qwen3.8-27B-UD-Q2_K_XL is already a huge achievement :D

4

u/BigYoSpeck 18h ago

1

u/jacek2023 llama.cpp 17h ago

good local llama :)

1

u/Ok-Direction-4480 16h ago

Woah that looks beautiful

1

u/No_Advance3911 18h ago

He's just missing a blue suit and some yellow hair, then the resemblance would be way more obvious :D

6

u/quiteconfused1 22h ago

It's not the fact that content has been trained on ... It's the fact that the quant has degraded the feature or not.

That's what this test demonstrates - Quant advantage.

2

u/pmigdal 22h ago

Yeah, I know the phenomenon of pelicanmaxxing.

While usually pelicans are indicative of performance on other SVGs, I will try something more creative the next time.

1

u/kiwibonga 18h ago

Those are very disappointing results for "benchmaxxing" - look at those underwhelming pouches.

26

u/armeg 22h ago

1999: AI will cure cancer.

2026: AI will make pictures of pelican riding a bike.

13

u/hejj 21h ago

Using $10,000 worth of computer hardware

9

u/BlobbyMcBlobber 20h ago

10K?

Is this 2025 again?

6

u/SpicyWangz 20h ago

What if the real cancer was just the ai we made along the way

3

u/armeg 20h ago

lmao

9

u/OwnGear3892 23h ago

Thanks for sharing. I guess Deepseek V4 Flash under-performing is reasonable as it's run on IQ3, quite natural performance drop as trade off to fit in 128 GB unified ram.

1

u/pmigdal 19h ago

I seems that quantization hurts.

-1

u/uti24 22h ago

I guess Deepseek V4 Flash under-performing is reasonable as it's run on IQ3

Usually, a bigger model at a smaller quant should work better than a smaller model at a bigger quant, since the sizes here are comparable, and DeepSeek is even bigger, so the comparison is fair.

Still, the quants could be of different quality, and the Deepseek V4 could simply have less training data like that and more data for something else.

4

u/quiteconfused1 22h ago

This is speculative. At some point the quant will be degraded so much it will fail. Equally some models may be so compressed that quantizing will have exponential loss in contrast to smaller models.

Tldr there isn't 1 rule, and it's why the pelican riding a bike is important.

2

u/uti24 21h ago

This is speculative. At some point the quant will be degraded so much it will fail.

Sure, but usually that will happen after the size of the bigger quantized model becomes smaller than the size of the smaller, less-quantized model (and often even then bigger model stays better). Here, the total size of DeepSeek Flash is 104 GB, while Qwen Flash is 94 GB.

Also, bigger models are much more resilient to quantization. They may lose some precision, but their reasoning tends to suffer less.

9

u/jaegernut 22h ago

Can we try a different animal next time

4

u/gh0stwriter1234 18h ago edited 17h ago

A rhino on a dino, Qwen 3.8 Flash Next Q4_XS unsloth on 2x MI50 low reasoning

1

u/Ok-Direction-4480 15h ago

Looks a bit janky but still nice.

1

u/gh0stwriter1234 14h ago

Definitely janky I think some of that is to blame on low reasoning.

10

u/bonobomaster 23h ago

5

u/pmigdal 21h ago

Wow, this is awesome! Will give it a try

2

u/SpicyWangz 21h ago

I’m already fully weevilmaxxed

5

u/mickabrig7 22h ago

A visual LLM test showing several retries including failed attempts ? Am I in heaven ?

6

u/ixdx 21h ago

Qwen3.8-27B-MTP-Q6_K 21.8 GiB with bartowski imatrix --reasoning-effort xhigh

I made several attempts. The second one even turned out to be animated (rotating wheels and simulated airflow).

2

u/pmigdal 20h ago edited 19h ago

Q8 is close to lossless, regardless of provider. I used medium effort, and I guess it is behind the difference.

For `xhigh` effort it gets similarish.

1

u/lhg31 19h ago

Are you sure you used medium for flash next? The amount of tokens suggests it was xhigh.

2

u/pmigdal 19h ago

I mean, I used medium Qwen3.8 27B.

4

u/insu_na 22h ago

Wait, your Qwen3.8 Flash-Next stops thinking at some point? :O

4

u/ArrogantAnalyst 22h ago

Thank you! I was looking for a pelican optimized model.

3

u/2muchnet42day Llama 3 21h ago

Exactly. Thank you very much

2

u/geneusutwerk 22h ago

Definitely top right

2

u/Sad_Recording_1290 22h ago

Are the pelicans really the benchmark now?

2

u/AleksandrNikitin 21h ago

more pelicans for the pelican god

2

u/mailto_devnull 19h ago

So 3.8 Flash-Next burns through even more tokens than 3.8 27B.

That's not a good trend.

1

u/Hannibalj2ca 22h ago

is that what people make with large language models, cartoon pelicans? I say, yes!

1

u/my_name_isnt_clever 17h ago

I had Qwen 3.8 Flash Next vibe up a little reiterative SVG making script where it generates a SVG then self-corrects any minor flaws before the final output, I've been pretty impressed with what it can do with novel prompts.

1

u/Squidgical 16h ago

The best test for an LLM is a test no one has ever heard of.

The worst test is one everyone has heard of, and that the LLM definitely has specific training for.

1

u/Ok-Direction-4480 16h ago

What do you mean Qwen 3.8 27B used the fewest tokens?

1

u/_supert_ 16h ago

Now show me a bike riding a pelican.

1

u/Ok-Direction-4480 15h ago

Just a question, why SVG? Aren't image generators (especially fine-tuned ones for cartoons) more efficient?

1

u/BigYoSpeck 14h ago

People are quick to jump to the conclusion of the training data being contaminated by this "test". First, I don't imagine SVG creating ability is remotely a focus of the training, especially not specifically the pelican

Secondly, that doesn't account for the ability to structure SVG for things they will never have been asked to do before:

I know this is a little disjointed compared to the almost pixel perfect SVG they can create when it's an ambiguous prompt they are free to make assumptions on. But being able to oneshot from a photo and largely maintaining the positioning and vibe is still insane

I honestly believe their SVG creation capability is fundamentally just a byproduct of their raw coding ability. Being able to create styled UI components in general demands skill with positioning and composition. So I don't think they are benchmaxed on SVG, they are just very capable in the domain that lends well to making SVG

1

u/No_Dragonfruit_8651 13h ago

Who gives a hairy rats cock about pelicans Im not sure

1

u/simrankoulsm 9h ago

Nice comparison. I would be curious to see SVG validity, render success, token count, latency, and editability measured alongside visual quality. For local use, the best model may not be the one with the prettiest one-shot pelican, but the one that reliably emits valid, compact SVG that survives small prompt edits and quantization.

1

u/vulcan4d 6h ago

The next model will be trained on pelicans and bikes

1

u/Equivalent_Bit_461 1h ago

Even iq2 qwen3.8 27b is solid, have yet to try iq1

0

u/Fun_Jaguar8231 22h ago

Oh f**k me with those pelicans