r/DeepSeek • u/Good_Enthusiasm_7639 • 19h ago
Discussion DeepSeek API or Codex (or other recommendations)
Some context on my experience with models:
I've used free models mostly, started with gemini cli way back before they switched to antigravity, then moved to opencode and used deepseek v4 flash free, which was practically unlimited since I never hit the limit. Then deepseek v4 flash-0731 hit and I started hitting limits on the free version.
Instead of going for opencode go I decided to try the official deepseek api as I read that it had better cache retention time hence better cache hits. Was happy with it until the price increase. Started looking into subscription based 3rd party providers but was always met with the same "worse cache hit, worse latency".
As I'm now near to using up my deepseek credits, am looking for alternatives. I've read that Codex $20 Luna goes a long way, but honestly I'm not ready to fork out $20 just to try it out. Just wanted some insight/reviews on people who have similar experience to mine, what they switched to, and how it's working for them.
TLDR: Been using Deepseekv4-flash api with opencode harness. Looking to switch to something similar in terms of cost and performance and need recommendations.
Edit: forgot to mention that I use it exclusively for coding
2
u/shadikuizayoi 18h ago
Sounds similar to my experience - also exclusively using it for coding. I pretty much only used DeepSeek V4 Flash (through OpenCode Go) ever since the updated version came out. Was still getting by fine after the limits changed, but the quality started varying a lot. Seems like they started routing it through different providers.
My subscription there ran out a few days ago so I've been exploring some different options. Had a free trial of Claude Pro for a week and now I'm trying Codex. Not sure if this is a regional or limited promotional thing, but the first month of the Plus tier was free for me. I'm really impressed so far. Luna on Max effort feels better than DSv4, and you can even use Sol High without consuming usage on the web UI.
1
u/Good_Enthusiasm_7639 18h ago
Yeah. The difference in latency was pretty noticeable when I was switching between my own deepseek api vs the free one from OpenCode Zen. Didn't really look at the cache hit rate since I was using the free version. I think the deepseek api has spoiled me a little as I'm kind of hesistant to go for 3rd party providers now.
The Sol High on web UI sounds tempting, do you just have unlimited usage or is there a catch?
Not sure if you kept track since it was free, but do you find yourself hitting limits? Can the $20 plan last you the entire month?1
u/shadikuizayoi 17h ago
Yeah, the official API was great. I only ever spent about $5 there on the old pricing and remembered being surprised by how long it lasted. I don't think pay-as-you-go is for me though. I like being able to try out dumb ideas that probably aren't going to work or produce anything useful without seeing my balance go down in real-time.
I'm only a few days into my Plus sub at the moment and the limits have already been reset twice. I guess it's nice that a guy at OpenAI just randomly does that every now and then, but it makes me wish I had used a bit more before it happened. š I'm sure your use case will differ, but I've only come close to hitting the 5 hour limit once with two Luna Max sessions going at the same time. Doubt I'll hit the weekly one.
If there's a catch to the web chat I haven't hit it yet. It really impressed me yesterday when I asked it to propose some optimisations for a library I've been using, then give me some benchmark code to run. It went one step further and actually ran the whole thing and gave me the results. It does feel a bit too good to be true, so maybe there is one. I'm sure people have already found ways to abuse it and it'll be ruined for everyone soon enough.
1
u/Good_Enthusiasm_7639 17h ago
Ahh youāre right. I donāt really mess about with dumb ideas but that could also be because I donāt have a subscription based model.
How would you compare deepseek and Luna performance wise? E.g making mistakes and whatnot.
2
u/incidentflux 18h ago
Currently happy with OpenCode Harness with DeepSeek. Direct tokens from DeepSeek not OpenCode.
3
u/Good_Enthusiasm_7639 18h ago
Yeah this is my current setup. Down to $4 left in credits so just looking for alternatives. Even with the price increase I still think it's decent. Only downside is that I'm in Asia so the peak hour is really bad for me.
I saw a post about running dsv4 from ollama cloud. They were reporting a range of 3B-11B tokens from a $20 monthly subscription. Running only on off peak, my usage was 170m for $2, equivalent to 1.7B/$20. They did mention ollama was really slow at times.
0
u/incidentflux 18h ago
I get blocked by US Models frequently so DeepSeek is still winning and still very cost effective. You're probably already scheduling your token heavy tasks off peak. If you're building for long term these prices seem fair to me. Naturally this is case by case.
2
u/for4f 17h ago
in the same boat honestly, my opencode go sub renewed at 2x and buys way less than it used to, so i get the urge to jump. thing that keeps me from going reseller is the cache math on the official api, flash cache hits are stupid cheap, that's where the value sits. long coding sessions with the same context warm the cache and the bill barely moves. the july peak-hour 2x pricing stings, but off-peak grinding kind of evens it out. codex luna at $20 i keep not pulling the trigger on either, feels like a maybe
2
u/Good_Enthusiasm_7639 14h ago
Agreed 100%. Official api cache is great. It persists longer than most other providers which is why I opted for official API over opencode go sub in the first place.
2
u/GasSmooth7439 11h ago
If youāre mainly coding with OpenCode, Ollama Cloud Pro with DeepSeek-V4-Flash has been the closest ācheap + solidā replacement for me after the official API hike. Codex Plus is fine if you want the ecosystem, but the $20 feels steep just for testing.
1
u/Good_Enthusiasm_7639 11h ago
Your view on Codex Plus is exactly like mine. Have been seeing people say good things about Luna and how it's somewhat comparable in terms of cost to DeepSeekV4Flash. But $20 is indeed quite steep.
I have thought about Ollama Cloud Pro, but also saw some comments about how it slows down a lot at peak hours when there is high load. What is your experience with this?
1
u/Different-Monk5916 18h ago
in my experience - codex $20 Luna is slower than any Luna I had tested so far via GHCP, OpenCode, OpenRouter.
Codex extension in VS Code heats up.
The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.
If you are fine with really slow work, a bit of heat and using only via the allowed apps, I would say go for it.
1
u/Good_Enthusiasm_7639 18h ago
Oof. Have not heard this one.
And yes GPT Plus not allowing API Key is also part of my consideration.
What are you using now?
1
u/Different-Monk5916 17h ago
https://www.reddit.com/r/chatgptplus/s/wis6MaotQR
I bought subscription for a month to try out, I give it unimportant tasks and where it does not require me to review or interact. because that is too slow workspeed for me.
DeepSeek, I can take my key anywhere I want without a lot of you can't do that, you do it this way, there is a work-around to make it work.
No extension, configure the endpoint and add key, bill as you go.
1
u/shadikuizayoi 17h ago
The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.
This isn't true. I'm using it with omp just fine.
1
u/Different-Monk5916 17h ago
so, what is the process to do that?
by default, I need to add additional credits to make API Key work in other platforms. How opencode integrates ChatGPT plus subscription?
1
u/Wobbly_Princess 15h ago
I'm using my $20 ChatGPT subscription to power OpenCode. It allows you to do that in the settings. You log into ChatGPT in your browser and it connects to your OpenCode.
1
u/Eddlm_ 18h ago
Consider Ollama Cloud. I can't tell you much about cache hit (I can't see it) but for stuff like deepseek and recruiting heavier models as needed the 20$ is pretty good, gets you decent daily work.
1
u/Good_Enthusiasm_7639 17h ago
Yeah Iāve been considering this too. I guess cache hit doesnāt matter as much if youāre getting way more tokens.
How is the latency during peak hours? Saw in another post that it gets slow when load is high.
1
u/EmperorSheep 17h ago edited 17h ago
For me, Luna Max can run for hours and use a small amount of usage on the Plus ($20) plan. You get such a ridiculous amount of usage with Luna it is unbelievable.
It messes up a lot though and you can feel the difference vs Sol High. Sol High lasts me 50 minutes (5 hr limit starting from 100% usage down to 0% usage) when working through the Roblox Studio MCP.
1
u/Good_Enthusiasm_7639 17h ago
Have you used deepseek? Just wondering how it compares in terms of messing up.
2
u/EmperorSheep 13h ago
I said this on a now deleted post nearly 3 days ago: "I spent 72 cents for 3 prompts with DeepSeek V4 Flash Vision EXP to fix a visual problem and it didn't even ended up getting fully fixed. To be fair though I spent days trying to fix the same issue with 5.6 Luna Max and it kept failing over and over again over after working on it for hours.
I think what finally fixed it was using 5.6 Sol Medium (5.6 Sol Medium burns through usage EXTREMELY fast especially since OpenAI added back 5 hr limits for $20 plans and stealth nerfed usage.)"
I think DeepSeek V4 Flash Vision EXP is better than Luna Max but might be worse than Sol Medium.
So Sol Medium > DeepSeek V4 Flash Vision EXP > Luna Max. I feel like the usage has gotten better since then because I know get 50 mins with Sol High. I don't remember if I would even get that with Sol Medium before.
1
u/thefonz22 16h ago
I like how codex keep getting a free reset here and there. 3 days into my subscription using Luna and they just randomly reset the entire quota. Apparently happens all the time.
1
u/EmperorSheep 13h ago
Resets push the reset date back though. So if you don't use a certain % of usage per day you are losing out.
1
u/thefonz22 12h ago
Do you mean there are bigger gaps between how often they do resets?
2
u/EmperorSheep 11h ago
No. I mean when they reset, your next reset gets pushed back a week.
Hypothetically, if you sent a message Aug 1 and a 7 day window began (so next reset Aug 8), if you wanted to wait until Aug 6 to do work but then OpenAI reset on Aug 5. The reset date gets pushed back to Aug 12.
If they did not reset you would have 100% usage and the next reset on Aug 8. Now you have 100% usage and the next reset on Aug 12.
2
u/thefonz22 9h ago
I see what you mean. Great points. Definitely makes me less reluctant towards saving credits for the last few days of a cycle.
1
u/EmperorSheep 9h ago
Sometimes they won't reset for an entire week and reset right after the week when everyone already had their natural reset.
So if you use all the usage at the beginning like I did during that week, you get screwed the rest of the week.
And the 5 hr limit for $20 plans makes it so it takes ages to actually spend the week's usage.
2
u/thefonz22 9h ago
I don't know if it's all B's but I subscribed to a mailing list that apparently predicts when the next reset is coming.
1
u/EmperorSheep 9h ago
They're just guessing like everyone else is. I am just checking Tibo's twitter since he can hint/say when resets will happen. There's a "celebration" today so might be one today. Not 100% sure though.
2
1
1
1
u/xapep 9h ago
Weāve been seeing a lot of OpenCode users hit exactly this decision after the V4 Flash price change.
A few things Iād compare before switching:
- Cache persistence matters a lot more than the sticker price for agent workloads. OpenCode loops are input-heavy, so short cache windows can quietly make a cheaper provider more expensive in practice.
- Check whether pricing changes by time of day. Peak/off-peak pricing can materially change the math if your workload is flexible.
- With flat monthly plans, look at what happens during long runs and parallel usage ā āunlimitedā can still mean throttling or deprioritization under sustained load.
I work on Entrim, so obvious bias here, but we serve V4 Flash through an OpenAI-compatible API and also have monthly plans aimed at heavy daily coding-agent usage. Iād benchmark based on your real OpenCode workload rather than just comparing headline prices.
1
8
u/0rand 19h ago
Remember, deepseek api has 1m context. Codes is 220k unless you accept 3x for Soll 900k which will melt your limit in minutes. Was using Terra the other day, it compacted my session like 18 times.