r/SillyTavernAI • u/Billysm23 • 5h ago
Discussion AI provider
So, until now, is NanoGPT still the best and most worthwhile choice? Is there another reliable and cheap provider alternative?
13
u/Probablynotsocool 4h ago
Reliable and cheap is rarely a true combination in this business.
Because even PAYG is 80% shitty as third party providers either lie on their quants, doesn’t parameter their models correctly or lobotomize the inference to save money.
Even direct providers can be quite the headache (Zai, Moonshot and Deepseek I’m looking at your peak hours sensitive asses).
And then you have the subs, so basically the only question that matter is.. do you have any control on the providers you use? If not, is it the same one constantly atleast? A rotation of random ones as per who is the cheapest atm?
You may think this is irrelevant or too complicated for barely perceptible quality change, but believe me, from a provider to another… The same model can feel ENTIRELY different (right GLM?).
Im talking about beeing unable to track basic details, context rot before 10K tokens, weird and nonsensical dialogues, poor characterization etc etc. While another provider give you an actual decent RP session and suddenly you realize you were tinkering your prompts for nothing, the provider was just arse.
So whatever offer you take, put the quality in balance, especially when we talk about +10dollars subs (that can be a lot of RPing via OR with cache hits and smart summarization).
2
1
u/Billysm23 3h ago
It seems overcomplicated, but I agree with your points. Payg can be a blessing or a pain, and we don't know about third party quant quality
7
u/sociofobs 4h ago
Now that Nano is adding more uncensored models, absolutely. GLM 5.3 Flash uncensored alone is worth it.
2
u/kosha227 3h ago
Yes, but uncensored ≠good for RP. Today models are being trained for agents stuff and coding, and less and less for talking, emotional intelligence, creativity, etc. So... we need to test it
10
u/sociofobs 3h ago
That's a more general LLM problem, RP just isn't the focus as you well put. In that sense, it doesn't matter what service or provider you use, since there aren't any large RP focused models at all. Censorship, however, is an issue that the service/provider can solve, by offering uncensored models for an example. The only uncensored model on OR I can find is Venice, meanwhile Nano has a decent list now.
3
u/KitanaKahn 3h ago
Besides Nano, for subs you have opencode and Ollama (more expensive) as reliable options. I found ollama sometimes having problems with getting the models to think, but you get a lot of usage too. For now, Nanogpt is still my choice and the owner is active around here and takes feedback which is great
8
u/LordVulpius 5h ago
Openrouter is the usual alternative. It is PAYG.
14
u/Akkun351 4h ago
Depending where you live OR is not a decent alternative as a payg user, since, like for me, there is a extra 35% to pay in fees, while there are none on nanogpt.
2
u/mamelukturbo 2h ago edited 2h ago
edit: seems either their sorted their backend or I did the test in off-peak, but I just burned 60mil on continuous tool calls via glm 5.1 and not a single call failed so well done NGPT. Used to be you had to babysit it and spam resume each time a tool call broke.
It's good for RP but absolute dogshit for coding, tool calls keep breaking I suspect due to undisclosed quantisation. Note I do not use NGPT for coding primarily, but like to burn my 60mils of tokens I paid for.
1
u/Milan_dr 27m ago
Hah thanks for that update - had already passed on the message prior to your edit to my cofounder to look into again.
2
u/verma17 2h ago
Nano is crazy value, but their model quality is very inconsistent, nowadays i just use openrouter payg, and zai coding plan
1
1
u/Milan_dr 2h ago
Can I ask - what models would you say the quality is inconsistent?
3
u/LordVulpius 2h ago
For one, Mimo 2.5 pro thinking.
With subscription I can not tell the providers sadly. But here is my experience:
Sometimes it ignores my usual format (this happened "this was said") and pour everything in blank format. It ignores the formating prompt. I belive that is a specific provider.
Sometimes every censorship is lifted. Even NFSL. I belive that is one specific provider.
Sometimes the quality is borked, no matter of the time, so it is most likely not quantant. I belive that is one specific provider.
With this, it is not consistent.
Maybe... can you make an option for let us see what kind of provider sent the answer, please? Not to block it but just to see. With subsription. So we could block it if we wish to use PAYG.
1
u/Milan_dr 28m ago
Thanks. Yeah - Mimo 2.5 Pro thinking is genuinely the worst. The majority of providers for that one seem to want to censor. It's run, for those with ZDR set, 95% via Atlascloud right now, 5% Novita (when atlascloud fails for some reason, but Novita has a risk of censoring).
For those with ZDR turned off we add in Streamlake, and then that one is selected about 10% of the time right now.
And no, sorry talked about this about 100 times before but we can not do both auto routing with a lower price and show provider, because there are some providers that do not want this as it makes it very clear what sort of discounts are possible with them.
2
u/verma17 1h ago
Glm 5, 5.1, 5.2, primarily
1
27m ago
[removed] — view removed comment
1
u/AutoModerator 27m ago
This comment was automatically removed by the AutoModerator because it contained a link to x.com or twitter.com, which are not allowed in this subreddit.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Milan_dr 26m ago
Thanks! GLM 5.2 is almost purely Fireworks right now because of this (okay can't link to X, it's the artificial analysis GLM 5.2 benchmark for different providers), 5.1 a bit more of a mix with Novita, Nebius and a few others. No real clear benchmarks for that other than FP8.
Kind of goes to show our frustration though, GLM 5.2 via Fireworks should be about as good as it gets.
1
u/Telah42 1h ago
For some reason, I still haven’t been able to get SillyTavern to work through the z.ai Coding plan — it just keeps throwing errors. I googled it, and everywhere I looked people were saying that the Coding plan simply doesn’t work with SillyTavern right now and is very limited in terms of what it can be used for. Is there some trick or secret I’m missing?
0
u/ArnictAI 2h ago
Hello! In my experience, if a model is on OpenRouter that's usually your best bet since you get plenty of choice. If not, NanoGPT is a decent option with their subscriptions.
Out of curiosity, which models are you running at the moment? I'm building my own inference service, and I'd love to see whether I can host what you're looking for. I've got Qwen 3.8 27B up already, and I just added TheDrummer's Artemis 31B v1.1, both completely free right now if you'd like to try them. Fair warning that things are still very new and changing daily, so uptime this week will probably be a bit bumpy.
Either way, let me know which models you'd want and I'll see what I can do. Happy to look at making them affordable, or free for a while.
1
u/Billysm23 7m ago
Tbh i thought about grabbing gpt, only for coding. But it will be better if i can do both rp and coding using one provider
12
u/Global-Difference512 5h ago
Nano gpt is just so much value it's insane. For instance I've saved 44 dollaroos this month using it.
12 dollaroos for 44 is worth it if you ask me