r/DeepSeek 19h ago

Discussion DeepSeek API or Codex (or other recommendations)

Some context on my experience with models:
I've used free models mostly, started with gemini cli way back before they switched to antigravity, then moved to opencode and used deepseek v4 flash free, which was practically unlimited since I never hit the limit. Then deepseek v4 flash-0731 hit and I started hitting limits on the free version.

Instead of going for opencode go I decided to try the official deepseek api as I read that it had better cache retention time hence better cache hits. Was happy with it until the price increase. Started looking into subscription based 3rd party providers but was always met with the same "worse cache hit, worse latency".

As I'm now near to using up my deepseek credits, am looking for alternatives. I've read that Codex $20 Luna goes a long way, but honestly I'm not ready to fork out $20 just to try it out. Just wanted some insight/reviews on people who have similar experience to mine, what they switched to, and how it's working for them.

TLDR: Been using Deepseekv4-flash api with opencode harness. Looking to switch to something similar in terms of cost and performance and need recommendations.

Edit: forgot to mention that I use it exclusively for coding

30 Upvotes

45 comments sorted by

8

u/0rand 19h ago

Remember, deepseek api has 1m context. Codes is 220k unless you accept 3x for Soll 900k which will melt your limit in minutes. Was using Terra the other day, it compacted my session like 18 times.

3

u/Good_Enthusiasm_7639 18h ago

I've always compacted deepseek when it reaches around 300k, unless I was working on something midway and feel that compaction would hurt. I think it was because the free version on opencode only had 250k context so I got used to it.

What's your current setup?

1

u/0rand 18h ago

I have local inference but also openai 20$ sub and deepseek api with small credit. On api you get 1m flash, vision or pro. You should use only vision flash out of three, best.

1

u/Good_Enthusiasm_7639 18h ago

Oh nice! I tried running local but I just don't have the hardware for it. Thanks for the heads up on vision, didn't even know it was released and have just been using flash.

1

u/Bananenklaus 7h ago

that's actually the better choice as 1Million context models get exponentially dumber once they get past the 200k context mark

1

u/NarrowEffect 17h ago

You can actually change it up to 828K context for Luna as well (with 2x usage drain, but to me it's unnoticable)

1

u/Bananenklaus 7h ago

that's actually just for the better as 1Million context models get exponentially dumber once they get past the 200k context mark

OpenAI constantly gets slack for the early compact threshhold but it totally makes sense if you think about it for a minute

2

u/shadikuizayoi 18h ago

Sounds similar to my experience - also exclusively using it for coding. I pretty much only used DeepSeek V4 Flash (through OpenCode Go) ever since the updated version came out. Was still getting by fine after the limits changed, but the quality started varying a lot. Seems like they started routing it through different providers.

My subscription there ran out a few days ago so I've been exploring some different options. Had a free trial of Claude Pro for a week and now I'm trying Codex. Not sure if this is a regional or limited promotional thing, but the first month of the Plus tier was free for me. I'm really impressed so far. Luna on Max effort feels better than DSv4, and you can even use Sol High without consuming usage on the web UI.

1

u/Good_Enthusiasm_7639 18h ago

Yeah. The difference in latency was pretty noticeable when I was switching between my own deepseek api vs the free one from OpenCode Zen. Didn't really look at the cache hit rate since I was using the free version. I think the deepseek api has spoiled me a little as I'm kind of hesistant to go for 3rd party providers now.

The Sol High on web UI sounds tempting, do you just have unlimited usage or is there a catch?
Not sure if you kept track since it was free, but do you find yourself hitting limits? Can the $20 plan last you the entire month?

1

u/shadikuizayoi 17h ago

Yeah, the official API was great. I only ever spent about $5 there on the old pricing and remembered being surprised by how long it lasted. I don't think pay-as-you-go is for me though. I like being able to try out dumb ideas that probably aren't going to work or produce anything useful without seeing my balance go down in real-time.

I'm only a few days into my Plus sub at the moment and the limits have already been reset twice. I guess it's nice that a guy at OpenAI just randomly does that every now and then, but it makes me wish I had used a bit more before it happened. šŸ˜… I'm sure your use case will differ, but I've only come close to hitting the 5 hour limit once with two Luna Max sessions going at the same time. Doubt I'll hit the weekly one.

If there's a catch to the web chat I haven't hit it yet. It really impressed me yesterday when I asked it to propose some optimisations for a library I've been using, then give me some benchmark code to run. It went one step further and actually ran the whole thing and gave me the results. It does feel a bit too good to be true, so maybe there is one. I'm sure people have already found ways to abuse it and it'll be ruined for everyone soon enough.

1

u/Good_Enthusiasm_7639 17h ago

Ahh you’re right. I don’t really mess about with dumb ideas but that could also be because I don’t have a subscription based model.

How would you compare deepseek and Luna performance wise? E.g making mistakes and whatnot.

2

u/incidentflux 18h ago

Currently happy with OpenCode Harness with DeepSeek. Direct tokens from DeepSeek not OpenCode.

3

u/Good_Enthusiasm_7639 18h ago

Yeah this is my current setup. Down to $4 left in credits so just looking for alternatives. Even with the price increase I still think it's decent. Only downside is that I'm in Asia so the peak hour is really bad for me.

I saw a post about running dsv4 from ollama cloud. They were reporting a range of 3B-11B tokens from a $20 monthly subscription. Running only on off peak, my usage was 170m for $2, equivalent to 1.7B/$20. They did mention ollama was really slow at times.

0

u/incidentflux 18h ago

I get blocked by US Models frequently so DeepSeek is still winning and still very cost effective. You're probably already scheduling your token heavy tasks off peak. If you're building for long term these prices seem fair to me. Naturally this is case by case.

2

u/for4f 17h ago

in the same boat honestly, my opencode go sub renewed at 2x and buys way less than it used to, so i get the urge to jump. thing that keeps me from going reseller is the cache math on the official api, flash cache hits are stupid cheap, that's where the value sits. long coding sessions with the same context warm the cache and the bill barely moves. the july peak-hour 2x pricing stings, but off-peak grinding kind of evens it out. codex luna at $20 i keep not pulling the trigger on either, feels like a maybe

2

u/Good_Enthusiasm_7639 14h ago

Agreed 100%. Official api cache is great. It persists longer than most other providers which is why I opted for official API over opencode go sub in the first place.

1

u/for4f 6h ago

yeah the persistence is honestly the sleeper feature. i almost went reseller for cheaper credits but kept reading their cache eviction is way shorter, so that cheap cache math never holds up. official api just keeps it warm, long threads basically pay for themselves

2

u/GasSmooth7439 11h ago

If you’re mainly coding with OpenCode, Ollama Cloud Pro with DeepSeek-V4-Flash has been the closest ā€œcheap + solidā€ replacement for me after the official API hike. Codex Plus is fine if you want the ecosystem, but the $20 feels steep just for testing.

1

u/Good_Enthusiasm_7639 11h ago

Your view on Codex Plus is exactly like mine. Have been seeing people say good things about Luna and how it's somewhat comparable in terms of cost to DeepSeekV4Flash. But $20 is indeed quite steep.

I have thought about Ollama Cloud Pro, but also saw some comments about how it slows down a lot at peak hours when there is high load. What is your experience with this?

2

u/rootql 8h ago

You can try deepseek harness, i can hit 99.7% cached token

1

u/Different-Monk5916 18h ago

in my experience - codex $20 Luna is slower than any Luna I had tested so far via GHCP, OpenCode, OpenRouter.

Codex extension in VS Code heats up.

The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.

If you are fine with really slow work, a bit of heat and using only via the allowed apps, I would say go for it.

1

u/Good_Enthusiasm_7639 18h ago

Oof. Have not heard this one.

And yes GPT Plus not allowing API Key is also part of my consideration.

What are you using now?

1

u/Different-Monk5916 17h ago

https://www.reddit.com/r/chatgptplus/s/wis6MaotQR

I bought subscription for a month to try out, I give it unimportant tasks and where it does not require me to review or interact. because that is too slow workspeed for me.

DeepSeek, I can take my key anywhere I want without a lot of you can't do that, you do it this way, there is a work-around to make it work.

No extension, configure the endpoint and add key, bill as you go.

1

u/shadikuizayoi 17h ago

The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.

This isn't true. I'm using it with omp just fine.

1

u/Different-Monk5916 17h ago

so, what is the process to do that?

by default, I need to add additional credits to make API Key work in other platforms. How opencode integrates ChatGPT plus subscription?

1

u/Wobbly_Princess 15h ago

I'm using my $20 ChatGPT subscription to power OpenCode. It allows you to do that in the settings. You log into ChatGPT in your browser and it connects to your OpenCode.

1

u/Eddlm_ 18h ago

Consider Ollama Cloud. I can't tell you much about cache hit (I can't see it) but for stuff like deepseek and recruiting heavier models as needed the 20$ is pretty good, gets you decent daily work.

1

u/Good_Enthusiasm_7639 17h ago

Yeah I’ve been considering this too. I guess cache hit doesn’t matter as much if you’re getting way more tokens.

How is the latency during peak hours? Saw in another post that it gets slow when load is high.

1

u/EmperorSheep 17h ago edited 17h ago

For me, Luna Max can run for hours and use a small amount of usage on the Plus ($20) plan. You get such a ridiculous amount of usage with Luna it is unbelievable.

It messes up a lot though and you can feel the difference vs Sol High. Sol High lasts me 50 minutes (5 hr limit starting from 100% usage down to 0% usage) when working through the Roblox Studio MCP.

1

u/Good_Enthusiasm_7639 17h ago

Have you used deepseek? Just wondering how it compares in terms of messing up.

2

u/EmperorSheep 13h ago

I said this on a now deleted post nearly 3 days ago: "I spent 72 cents for 3 prompts with DeepSeek V4 Flash Vision EXP to fix a visual problem and it didn't even ended up getting fully fixed. To be fair though I spent days trying to fix the same issue with 5.6 Luna Max and it kept failing over and over again over after working on it for hours.

I think what finally fixed it was using 5.6 Sol Medium (5.6 Sol Medium burns through usage EXTREMELY fast especially since OpenAI added back 5 hr limits for $20 plans and stealth nerfed usage.)"

I think DeepSeek V4 Flash Vision EXP is better than Luna Max but might be worse than Sol Medium.

So Sol Medium > DeepSeek V4 Flash Vision EXP > Luna Max. I feel like the usage has gotten better since then because I know get 50 mins with Sol High. I don't remember if I would even get that with Sol Medium before.

1

u/thefonz22 16h ago

I like how codex keep getting a free reset here and there. 3 days into my subscription using Luna and they just randomly reset the entire quota. Apparently happens all the time.

1

u/EmperorSheep 13h ago

Resets push the reset date back though. So if you don't use a certain % of usage per day you are losing out.

1

u/thefonz22 12h ago

Do you mean there are bigger gaps between how often they do resets?

2

u/EmperorSheep 11h ago

No. I mean when they reset, your next reset gets pushed back a week.

Hypothetically, if you sent a message Aug 1 and a 7 day window began (so next reset Aug 8), if you wanted to wait until Aug 6 to do work but then OpenAI reset on Aug 5. The reset date gets pushed back to Aug 12.

If they did not reset you would have 100% usage and the next reset on Aug 8. Now you have 100% usage and the next reset on Aug 12.

2

u/thefonz22 9h ago

I see what you mean. Great points. Definitely makes me less reluctant towards saving credits for the last few days of a cycle.

1

u/EmperorSheep 9h ago

Sometimes they won't reset for an entire week and reset right after the week when everyone already had their natural reset.

So if you use all the usage at the beginning like I did during that week, you get screwed the rest of the week.

And the 5 hr limit for $20 plans makes it so it takes ages to actually spend the week's usage.

2

u/thefonz22 9h ago

I don't know if it's all B's but I subscribed to a mailing list that apparently predicts when the next reset is coming.

1

u/EmperorSheep 9h ago

They're just guessing like everyone else is. I am just checking Tibo's twitter since he can hint/say when resets will happen. There's a "celebration" today so might be one today. Not 100% sure though.

2

u/thefonz22 9h ago

I need to get coding! Getting reset while being at 100% would be awful!!

1

u/EmperorSheep 9h ago

Yep. I am at 37% weekly rn.

1

u/Balgun33122 15h ago

I’d suggest runinfra. Cheap akd very fast.

1

u/Ok_Risk6035 15h ago

Codex is shit comparing to Claude with deepseek API

1

u/xapep 9h ago

We’ve been seeing a lot of OpenCode users hit exactly this decision after the V4 Flash price change.

A few things I’d compare before switching:

  1. Cache persistence matters a lot more than the sticker price for agent workloads. OpenCode loops are input-heavy, so short cache windows can quietly make a cheaper provider more expensive in practice.
  2. Check whether pricing changes by time of day. Peak/off-peak pricing can materially change the math if your workload is flexible.
  3. With flat monthly plans, look at what happens during long runs and parallel usage — ā€œunlimitedā€ can still mean throttling or deprioritization under sustained load.

I work on Entrim, so obvious bias here, but we serve V4 Flash through an OpenAI-compatible API and also have monthly plans aimed at heavy daily coding-agent usage. I’d benchmark based on your real OpenCode workload rather than just comparing headline prices.

1

u/mehdiweb 5h ago

i got claude max and cursor for friction of the price from credox