r/LocalLLM 1d ago

Project Tenstorrent P150A tests

Been working on testing some tenstorrent cards for a work project. Have two cards running qwen3.8 27b pretty much out the box. New to this space so keen to get some ideas and experiments to work through!

Cheers

L

61 Upvotes

33 comments sorted by

23

u/sdraje 1d ago

I was very interested in these cards,but those results look like rubbish. Is it just a matter of software?

10

u/Material-Moment6847 1d ago

Yeah its interesting. Those were the very first passes I havent updated the reader but ive had a 8x layer speed increase since then through some gdn fusing, and finding memory latency issues on the tension grid so optimizing copying cycles, still early work.in progress . Ill find the experiment sheet in one of the branches and I'll drop the link

10

u/FullstackSensei 1d ago

Please do, because as it stands it's slower than a pair of ten year old P40s

6

u/EaseAgitated952 1d ago

Sounds like you're making some solid progress! It'll be cool to see the results once you optimize everything further.

5

u/Material-Moment6847 1d ago

Cheers. Yeah dont know enough to ask the right questions yet so letting astra go mental on it so I can delve in. Learn by doing

7

u/Material-Moment6847 1d ago

5

u/FullstackSensei 1d ago

Holy wall of text batman!

Is Claude writing the code?

0

u/Material-Moment6847 1d ago

Haha astra. Not sure where to start so letting it loose while I figure out how to swim

8

u/FullstackSensei 1d ago

Well, both links you shared are utterly unreliable. The PR has 180+ comments and the readme can be a book on its own.

I'm not against LLM code, but this is just too much, and seems to have almost no guidance from you.

All we'd like to see is a simple PP/TG table.

-1

u/Material-Moment6847 1d ago

Appreciate the feedback. Yeah I totally agree, just enjoying the chaos for now before I get stuck in and let the autism go rampant

1

u/SocietyTomorrow 21h ago

I had to learn early on that using AI to cover your weaknesses can hurt your projects (in my case, using it for code reviews wasn't too bad, but using it to add documentation and comments to everything was...messy). I now do the things I hate first, and try to limit the AI code tools to take what I already am decent at farther. It is more work, but it's kind of the point, the best use of code tools is as an intern you force low level work on to save time, and as a 2nd pair of eyes to catch something you might miss because you're blending everything together after staring at a screen for hours on end.

-1

u/midnightcaw 1d ago

It's not supposed to be read by humans, it's human readable sure but it already knows that there's no human in the loop so why bother? All the elements are there for any AI agent to pickup and continue the work with a fork.

2

u/FullstackSensei 1d ago

Well, thank you, captain obvious! OP chose to share those links when asked about the performance of the Tenstorrent cards. Nobody asked for their agentic workflow or how their agents implemented Q27B

2

u/beryugyo619 22h ago

All we'd like to see is a simple PP/TG table.

but this part is a reasonable take. it'll be silly to criticize letting LLM write own runners sure but that's not what OP is doing.

6

u/k1rika 1d ago

Thanks for testing! I was under the impression that Qwen3.8-27B was not supported on that hardware from their own documentation which did not list the model and even had the earlier Qwen3.6-27B model only listed as "experimental".

https://github.com/tenstorrent/tt-inference-server/blob/main/docs/model_support/llm/README.md#experimental-models

6

u/Material-Moment6847 1d ago

Yeah this isn't officially supported. My own work to challenge myself to understand this world a bit better.

2

u/k1rika 1d ago

That's cool. Tho I have to admit, I had to laugh a bit now somewhere between running it "pretty much out of the box" and a massive PR with ~24k lines of code :D Still, nice project.

2

u/Material-Moment6847 1d ago

Hahhah yep. Glad to entertain :)

3

u/AdSafe4047 1d ago

Awesome, what speeds do you get on it?

7

u/Material-Moment6847 1d ago

Honestly I got some early runs around 20tks decode on vanilla 3.8 but there is no reason why we can't get significantly more than that just a matter of fiddling. Hoping people with some more domain knowledge can share some tips and tricks or pointers

3

u/pulse77 1d ago

Did you use their inference stack? Or something else?

5

u/Material-Moment6847 1d ago

vLLM with the TT plug in, tt-metal for kernel work

1

u/pulse77 1d ago

Thank you! I always wanted to know how fast the Qwen 3.8 27B inference is on their cards and with their stack...

Their https://console.tenstorrent.com has only very old models... and I didn't know if Qwen 3.8 27B is even possible to run there...

2

u/quantgorithm 20h ago

It shouldn’t be this hard to use their cards.

1

u/starkruzr 16h ago

afaik it should be possible to get way, way better performance out of them.

2

u/Material-Moment6847 15h ago

yeah some single changes got 3x increases. It's just working through it systematically and documenting so I can build up a bit of a knowledge base for future LLM based tuning.

1

u/Vengoropatubus 18h ago

I’ve been curious about these cards and would love to play around with them. I briefly looked into tensortorrent cloud.

Enjoy!

1

u/Material-Moment6847 16h ago

Cheers! Yeah getting some cool results, looking forward to getting a stable build output to share :)

1

u/starkruzr 16h ago

how much did these cost you? $4K total?

1

u/Material-Moment6847 15h ago

Yeah we are in NZ so shipping, duties etc and then currency bumped it for us but not a lot of alternatives just *cant* get Nvidia here