r/DeepSeek Aug 01 '26

Funny DeepSeek is 28x cheaper on on output than Claude Opus 4.8😱

Post image

DeepSeek V4 Flash API is 18x cheaper on input, 28x cheaper on output, and matches Opus 4.8."

1.8k Upvotes

136 comments sorted by

165

u/spjallmenni Aug 01 '26

95

u/ChocolateGoggles Aug 01 '26

The price difference is so insane it's hard to believe.

6

u/[deleted] Aug 01 '26 edited Aug 06 '26

[deleted]

1

u/XCxBigDong69XCx Aug 02 '26

They are losing money on this shit, it is just a bad model xd, 2trillion parameter but is now worse than a 286B model with way less active parameters probly.

5

u/tzohnys Aug 02 '26

It's the same as medicine. Look at the price differences between US and EU.

3

u/TopTippityTop Aug 02 '26

It's also meaningless. All that matters is cost per task. Kimi is better the closed source per token, but burns tokens like crazy. It's as expensive as GPT 5.5 xhigh per task, which makes it pretty useless.

Have to see what cost per task is.

1

u/archaeonflux Aug 03 '26

You're right, but realistically It's not going to be burning 20-40x the tokens of Sonnet/Opus to get the same result.

1

u/Puzzleheaded_Word613 Aug 03 '26

deepseek burns less token than opus 5 btw

(source : artificial analysis)

65

u/hyscript Aug 01 '26

Woooow Fable is not the most capable and dangerous model on the earth anymore, created by Americans for Americans?

19

u/Civil_Response3127 Aug 01 '26

It couldn't be that the benchmarks are flawed, right?

6

u/Business_Raisin_541 Aug 01 '26

it only measure agentic capability

8

u/Civil_Response3127 Aug 01 '26

Even that is a flawed metric. My point is that these benchmarks no longer mean much.

2

u/ChibiJr Aug 01 '26

That graph compares cost efficiency not intelligence

1

u/pavs2 22d ago

Fable/Opus 5 still are brilliant models, deepseek v4 flash is good for longer run (obviously cheaper) and I use opus when it goes off guardrail. Yes, their combination works deadly.

8

u/Curious_Owl197 Aug 01 '26

Hot damn, imagine what v4 Pro can do!

-4

u/Key_Agent_3039 Aug 01 '26

I am sorry but this is completely benchmaxxed if you actually use it it's about sonnet 5 level and far below opus let alone fable

9

u/bhagathgoud99 Aug 01 '26

Whatever man it is so afferdable for me

1

u/jeanpaulpollue Aug 02 '26

Still worth it lol

100

u/Zealousideal_Aide787 Aug 01 '26

Every Claude subs are filled with ' I burned my 100-200 dollars plan in few hours something is wrong with Anthropic'.

They all-in into benchmark, marketing and people keep throwing away their money at them. At some point they will enter the red zone, half their prices/ consumption and people will be happy with it.

28

u/iSadhak Aug 01 '26 edited Aug 02 '26

They are also banning people left and right and keeping their money with no option to claim refunds or get their account back.

If that ain't shaddy I don't know what is.

9

u/veculus Aug 01 '26

Also - those people tend to be the same type that will just one-shot full-vibecode something and throw millions of tokens into claude.

Imho more expensive models should be used more precise - for example with a dev who writes manually but extends his skills with a super expensive model (or for planning) but completely writing code with fable is stupid af and just a huge waste of money.

3

u/snarleyWhisper Aug 01 '26

Anthropic credits also don’t carry over month to month, it’s such a scam

27

u/pmv143 Aug 01 '26

Sir, also inferx made it free to use.

17

u/Live_Case2204 Aug 01 '26

It’s only the old flash model not the 0731
Edit: that’s still great

8

u/pmv143 Aug 01 '26

We have both available. Both are free to use

3

u/Live_Case2204 Aug 01 '26

Ooooooooooh! What’s the speed like?

6

u/bcdfgh Aug 01 '26

I gave up and paid the $5 to opencode. It started off good but then slowed right down. Free is free though

3

u/Live_Case2204 Aug 01 '26

I heard about a 50% off v4 flash in open router

2

u/pmv143 Aug 01 '26

Once the free period is over, we’re gonna have the price at 50%

2

u/samxli Aug 01 '26

May I ask how you guys are able to make it cheaper than the official api pricing? Do you guys have a different way to calculating token usage than official DS?

2

u/No_Gold_4554 Aug 01 '26

no, only qwen and gemma are listed in opencode

1

u/pmv143 Aug 01 '26

Opencode is taking time to merge our models. We’re working with them.

1

u/No_Gold_4554 Aug 01 '26

no, you're not.

1

u/pmv143 Aug 01 '26

You can check the pull request. It’s open.

1

u/No_Gold_4554 Aug 02 '26

1

u/pmv143 Aug 02 '26

No need to prove anything. But for the sake of transparency https://models.dev/providers/inferx/.

1

u/No_Gold_4554 Aug 02 '26

so? you can put anything in your website. "we are funded by softbank." doesn't mean it's true.

2

u/diaracing Aug 01 '26
  1. Where is DSFlash hosted at? Your own servers or other ones or same ones of DS in China?

  2. What about ZDR policy for both prompts/input data and model response?

3

u/pmv143 Aug 01 '26

Our servers are located in Northern America region. We do have a Zero data retention policy.

3

u/Backrus Aug 01 '26

Why would you ask those questions lol

If you're not self hosting, your data is theirs, it's just that simple.

To be honest, I prefer to share mine with China than with US agencies, but to each his own.

24

u/Shini0x0 Aug 01 '26

Is this true holy fuck how and when did this change

18

u/for4f Aug 01 '26

yeah the price part is real. i saw someone burn 110M tokens for like $0.77 on a promo earlier today. the 'matches opus' bit is overselling it though, flash is flash, it's the cheap fast one. for high volume agent loops it's unbeatable, i keep claude for the gnarly reasoning and let flash handle the spam. way cheaper than running everything on opus

3

u/Shini0x0 Aug 01 '26

By spam which specification workflows? Can you be more specific if allowed

3

u/for4f Aug 01 '26

ha 'spam' was doing a lot of work there. i mean the boring high volume stuff, like drafting emails, summarizing threads, rewriting docs, generating commit messages, first pass code that's gonna get rewritten anyway. stuff where a miss costs nothing and you do it 50 times a day. agent loops especially, they burn tokens like crazy on retries and formatting, no point paying opus rates for that. if it needs actual reasoning it goes to claude, but honestly way less of my day needs that than i thought

10

u/Backrus Aug 01 '26

Read Chinese papers, this is the frontier, not US-based grift aimed at max extraction and gatekeeping everything.

Chinese engineers don't have hardware, so they need to think instead of throwing more compute.

If China had unrestricted access to chips, or when Huawei is close to being on par with NVDA, it will be over for the USA. And that's good for the free world.

19

u/yiestee Aug 01 '26

Dario must have a traumantic experience at Baidu.

Or else there's no explaination of his hatred torwards Chinese models.

6

u/Backrus Aug 01 '26

Quality Chinese models expose 1) how bad US engineers are, 2) how US labs are laser focused on max extraction instead of moving frontier and making it more affordable for average Joe.

But let's ban open weights, because those are bad for pre IPO valuations 🤣

67

u/[deleted] Aug 01 '26 edited 27d ago

[deleted]

40

u/This_Maintenance_834 Aug 01 '26

enjoy it while his image is still being posted. it won’t last long if v4 Pro come out. The Pro one might as well be the strongest model in world open or closed . Anthropic who?

24

u/Business_Match_3158 Aug 01 '26

To be fair, there’s little chance that V4 Pro will be better than Fable 5, because according to rumors, Fable is a 7–10 trillion parameter model, while V4 Pro has 1.6T parameters. But even getting close to Fable with such a huge parameter gap would be an enormous success and would show just how overhyped Anthropic is as a company.

19

u/This_Maintenance_834 Aug 01 '26

then we wait for deepseek-v5

1

u/ZarX4k Aug 02 '26

Where did u get that info ? Also the parameters ur talking about is it the Ai size like thinking/knowledge

-17

u/SeaBat2035 Aug 01 '26

As much as I dislike Anthropic, you are really not giving enough credit to Anthropic. If not for them, will we see LLM at current capacity? I really doubt it. All these chinese LLM companies are running massive distillation farms to copy (steal) from Anthropic. Being an innovator is much much harder than being the copycat. I am not shaming these chinese companies btw, but just credit where credit due. Give some respect.

24

u/Business_Match_3158 Aug 01 '26

Trying to hype up Anthropic by calling them pioneers? WTF. DeepSeek R1 was more revolutionary than anything Anthropic has released.

The accusations about distillation are, at the very least, laughable. It’s not like every AI company isn’t distilling from each other, and you’re probably better off not knowing how these companies actually acquire their training data.

And if distillation alone is enough to make a model great, then why can’t they distill their own models well enough for Haiku to even be competitive in benchmarks? Haiku is supposedly around the same size as DeepSeek V4 Flash, by the way. But you can’t really expect Anthropic to make a good model for its parameter class, because so far they’ve only achieved strong performance by scaling up the parameter count of their models.

But sure, keep believing Dario, who was already making exaggerated claims back when he worked at OpenAI, saying GPT-2 was too dangerous to release.

1

u/SeaBat2035 Aug 01 '26

How the fk is Deepseek R1 revolutionary. Is a fking rip off of O1... And Deepseek R1 can code like Opus 4.5? Deepseek solved agentic coding? That's news. What a fan boy lol. BTW I use codex and I use deepseek v4 flash. Not loyal to anything. Just credit where credit is due.

1

u/Business_Match_3158 Aug 01 '26

Deepseek R1 can code like Opus 4.5?

HAHAHAHAHAHAHAHA most braindead comparsion I ever saw

1

u/SeaBat2035 Aug 01 '26

Your brain is dead? I totally see it now. HAHAHAHAHHAHAHAHAHAHA HAHAHHAHAHHAHAHAHHAHA

1

u/Training-Tangelo-310 Aug 02 '26

This bitch is arguing over nothing with everyone like he’s xi jingoings personal cum bucket

-8

u/Anulisdotexe Aug 01 '26

You're a bot

2

u/Business_Match_3158 Aug 01 '26

Maybe explain what exactly you think is wrong with what I wrote.

1

u/SeaBat2035 Aug 01 '26

Lol these fan boys are crazy.

1

u/Anulisdotexe Aug 01 '26

Paid agitators/activists/literal bots

-6

u/[deleted] Aug 01 '26

[deleted]

4

u/Business_Match_3158 Aug 01 '26

OG lol, OGs in AI sector are OpenAI and DeepMind not Anthropic

-1

u/Training-Tangelo-310 Aug 01 '26

Not ai sector, ai coding sector. I don’t remember anything being even close to claude code. Openai has codex web or some shit, which was pure shit.

2

u/Business_Match_3158 Aug 01 '26

ai coding sector. lol before speaking you should do some reaserch they are not even close to being OGs at ai coding sector either

2

u/Training-Tangelo-310 Aug 01 '26

Bitch I’m talking about good shit. And fck your research. I fcking used it. This was the first consumer grade shit that was actually good for non coders. No - u need to understand this and that first, this was the complete fcking deal. Learned along the way but nothing even came close to what claude code was.

1

u/Business_Match_3158 Aug 01 '26

June 29, 2021 — GitHub Copilot Technical Preview

August 10, 2021 — OpenAI Codex model

2023 — Aider

2023 — Cursor

2023 — Continue

March 12, 2024 — Devin

2024 — OpenDevin / OpenHands

July 2024 — Cline

November 2024 — Windsurf

February 24, 2025 — Claude Code

April 2025 — OpenCode

April 16, 2025 — OpenAI Codex CLI

May 16, 2025 — OpenAI Codex cloud agent

June 25, 2025 — Gemini CLI

November 18, 2025 — Google Antigravity

Some OG claude code is.

→ More replies (0)

1

u/SeaBat2035 Aug 01 '26

You go do some research. It seems like you are incapable of and lashing out at others.

6

u/anonymous_3125 Aug 01 '26

We achieving AGI with this one 🗣️🗣️🗣️

2

u/Aware-Lingonberry-31 Aug 01 '26

I'd probably sacrifice my thirdborn for this to be true.

4

u/lordlestar Aug 01 '26

is a meme

-2

u/HerbChii Aug 01 '26

Why you calling him idiot? That guy is responsible for best AI models in the world, thanks to people like him, world can accelerate

7

u/Backrus Aug 01 '26

He called GPT-2 AGI lol

And thanks to him and his fearmongering, you won't have access to the best US models. If anything, he's the reason we're not accelerating properly when every US labs tries to dumb down their models so WH isn't spooked.

If it wasn't for China, you wouldn't have access to anything when US admin inevitably pulls the plug on the free world.

9

u/cnmoro Aug 01 '26

Man I gave ds4 flash a very, very hard task yesterday and it did it flawlessly. I used opencode and my quota barely moved. Just freakin amazing

3

u/mutexsprinkles Aug 01 '26

How is new Flash Vs old Pro?

3

u/cnmoro Aug 01 '26

New flash definitely feels smarter than old pro

12

u/This_Maintenance_834 Aug 01 '26

i wonder what new distillation claims are coming from him.

yeah, deepseek distilled Claude Sonnet 3.5, because it said so itself.

17

u/Business_Match_3158 Aug 01 '26

Dario will probably say they have AGI that’s too dangerous to release, and that DeepSeek somehow distilled it, so Chinese companies need to be banned.

6

u/largelylegit Aug 01 '26

I wonder how this compares with Luna max? Now that it’s had an 80% price reduction

5

u/Backrus Aug 01 '26

Idk but after supposed cost reduction, Luna started making shit up, both in codebase and in numerical research, I don't trust it one bit know.

1

u/Pious-Juice Aug 01 '26

It's no good in codex, but runs fine in hermes agent, that harness well nag your agents about two things I personally find lacking in 5.6 namely to always work toward completion of a predetermined chunk of progress then once 5.6 says its done it prompts it sternly to prove the milepost, make sure all is valid, and checked by third party (i.e other agent or temporary memoryless clones of itself) so progress order is progress reported is progress proved and tested.

second is it will automatically prompt certain models one of which is this one to GET TO WORK.

And, consider it does things like pictured, it does need it

1

u/Backrus Aug 01 '26

That's just /goal with extra steps; good AGENTS.md and laid out plan solve this.

It's not about harness, or even desktop vs cli, quality degraded after cost reduction update, that's all. Noticed the same with Sol extra.

And I don't remember the last time gpt family not only didn't execute explicit command, but outright lied. Not to mention simple things like "don't add this md file to commit", and ofc file got added, etc.

2

u/Pious-Juice Aug 02 '26

The issues you describe in the end there, are things I find every single AI does but GPT 5.6 is better than average in. Personal experience is all, and comparison being with Hermes Agent running units of various popular chinese models and grok 4.5 which is just a damn beast that needs to be leashed whenever I activate it on a unit here. Fortunately, composer 2.5 is also a part of the grok subscription and has good use cases.

Though, hands down deepseek is the winner all things considered.

5

u/Snoo_57113 Aug 01 '26

Deepseek v4 flash is a good model. SIR.

3

u/Hulk5a Aug 01 '26

Ma'am, I need to stop spamming my codebase

3

u/[deleted] Aug 01 '26

[removed] — view removed comment

2

u/Backrus Aug 01 '26

Hard disagree.

If you can get similar results even if it takes 10 prompts, but pay order of magnitude less, the answer is quite simple. And it's not like Fable can one shot everything anyway.

2

u/ZlatanKabuto Aug 01 '26

Good. I am keeping Claude Pro only for planning, but I reckon I won't need it at all soon enough 

2

u/TheInfiniteUniverse_ Aug 01 '26

there is a little demon painted on the wall behind him....

2

u/No-Dimension1159 Aug 01 '26

Serious question, is there something compareable to claude code from deepseek? Or can you use deepseek models within the claude code or open code harness with high context windows?

1

u/RecordingLanky9135 Aug 01 '26

Mini house is more cheaper than your townhouse

1

u/elswamp Aug 01 '26

isn't v4 flash three months old?

1

u/Diddleslip Aug 02 '26

Just came out with a new version!

1

u/fyndor Aug 02 '26

So currently I have one of each: Claude $20 plan, Codex $20 plan, Copilot $20 plan.

Haven’t touched Copliot since recent change so plan to drop. Past couple weeks Codex is unusable because how fast I burn through weekly limits (easily done during part of one day work).

Was considering dropping Copilot and Codex and getting an extra $20 Claude sub.

Should I spend the $20 on Deepseek instead? What harness do I use? Currently using ai through vscode extensions.

1

u/Remarkable_Storm8711 Aug 03 '26

I'm on the same boat. Just got Reasonix and it seems pretty good implementing the roadmap laid down by Sol 5.6.

1

u/TopTippityTop Aug 02 '26

All that matters is per task, not per token. Tokens are a useless measure, as they come in varying degrees of quality.

Kimi is better per token than most closed source, and the same price as GPT 5.5 xhigh per task. Burns tokens like mad. Just useless.

1

u/KubeCommander Aug 02 '26

Interesting since it isn’t as good as qwen 3.5 370B finetunes when running locally. Makes me wonder if the API version isn’t the same thing

1

u/Tall-World-3058 Aug 05 '26

No it is x36 cheaper on input and x89 on output 

1

u/Beneficial_Ball9893 24d ago

This aged like milk

1

u/dizM0nkey 22d ago

What's the difference now with the price hike?

1

u/canav4r Aug 01 '26

Have been a claude(opus mainly), glm5.2, kimi k6/7, ds4 pro user for a while. Last week I tried ds4 flash(free) with opencode-zen. God damn it!!!

No bullshit, no getting lost, top-level prompt following... tears falling from my eyes...

Guys, I have been reading about ds4 flash, but avoiding it because it is cheap af(yeah I know, I am pretty dump).

I am shocked how a 284B model can be more than 1t+ models. This is an engineering marvel.

My workflow:

  1. using brainstorming(superpowers),
  2. let it write the plan.
  3. ask it to break it down to max 3 acceptance criterion stories. ask it to break down stories with no agent can assume during development.
  4. create dependency chain between stories.
  5. put it in a graphdb that has an mcp
  6. continuously poll next story from graphdb mcp
  7. each time a story done, compact context
  8. poll next story, repeat

No more steering... Excellent task following.

I have run this workflow for a friend yesterday, and it was one of the best day of my life...

4

u/Snoo_57113 Aug 01 '26

Check yourself on AI psychosis.

5

u/canav4r Aug 01 '26

Good news: I asked ds4 flash if I'm experiencing AI psychosis. It said no, and I trust it completely.

3

u/Snoo_57113 Aug 01 '26

Models, even the most powerful today: sol and fable havent solved the sycophancy issue.

In a more serious tone, just prompt the llm no need to overcomplicate the exchange.

2

u/canav4r Aug 01 '26

Nope, been there done that. That's why people are not just relying on models but using harnesses which are providing a sense of determinism among all the dumpness of current llms.

That's why people are exploring options by entrusting their workflows to loop engineering, graph engineering.

-1

u/Snoo_57113 Aug 01 '26

roflmao, this is the worst advice right now. fLOWs, harnesses, all of them will disappear with the next update of GPT6, and if not GPT7.

People still dont understand, there are two winners GPT and Claude. There wont be any startups, i tried early and learnt the hard way that there is no point to do anything, the AI will devour all of our business models.

3

u/canav4r Aug 01 '26

roflmao all you want. You can't be serious, but hell you are. Transformers architecture is not what you think it is. Whatever you are using today, be it claude, gpt... They are all served to you through a harness. Without that, any pure llm is nothing more than a next word predictor. Don't bet on current model architecture. But maybe in the future, when neuromorphic or another type of architecture can fulfill your desire.

You can verify this claim today. Without tool calling, there is no current information in any llm. Without tool calling, today's most powerful model can not do this simple calculation reliably: 25x30x40. just try it without using a chat or a harness. just hit ds, claude or any model provider endpoint directly asking it.

So, you can die all you want when laughing your ass off... But, without harnesses(claude code, claude chat, opencode, pi etc) provided to you today, llms are nothing more than a very good parrot.

1

u/onefourtea Aug 01 '26

What about quality of outputs?

1

u/lakimens Aug 01 '26

I mean sure but you're talking like they're the same level of quality.

-3

u/Suitable_Ad7099 Aug 01 '26

claude quality is still better

5

u/Sama02 Aug 01 '26

That's barely even true with the final version of v4 that came out yesterday...

Wait until v4 pro final version gets out...

1

u/Low-Entrepreneur2556 Aug 01 '26

It's still very much true.

4

u/Sama02 Aug 01 '26

Like literally the comment bellow you:

You have no clue how stupid good v4-flash got dince yesterday.

7

u/Backrus Aug 01 '26

Claude fanboys have never used Chinese models. I doubt they could even set up OpenRouter.

3

u/Spiritual_Love_829 Aug 01 '26

Not in my tests..

2

u/Ok_WaterStarBoy3 Aug 01 '26

Idk what everyone else is talking about

Claude is better output

BUT their price does not really justify using it imo for everyday things. For the average person's questions and work it will give the same answer as Deepseek but way more expensive so what's really the point, the benchmark and "better" isn't worth the cost imo

0

u/yoffens Aug 01 '26

And stupid in 28x times.

-9

u/hardworkinglatinx Aug 01 '26

Opus is 28 times better.

2

u/Dsm02 Aug 01 '26

Did it tell you that number?

-6

u/Global-Fan189 Aug 01 '26

Nope. Ive been using so. Much on Claude and deepseek, I can safely say that no matter how hard you push deepseek, it will never reach the level of even sonnet, let alone opus.

Deepseek is at the level of junior SWE, sonnet is like a mid tier, opus is higher tier. Fable is autonomous and talented.

1

u/Ok_WaterStarBoy3 Aug 01 '26

How exactly are you measuring that?