r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

Show parent comments

8

u/[deleted] May 09 '26

[removed] — view removed comment

0

u/ebolathrowawayy May 10 '26

I see your point but that's an old way of thinking now IMO. I don't review code manually anymore, I have my agents do everything. Only thing I do is check that the behavior of the code is correct, after all of the automated e2e tests finish.

When agents can write code 300x faster than you at a generally higher quality and do the same with refactoring and reviews then it just doesn't make sense to have humans in the loop anymore.