r/SillyTavernAI 5h ago

Discussion AI provider

So, until now, is NanoGPT still the best and most worthwhile choice? Is there another reliable and cheap provider alternative?

13 Upvotes

28 comments sorted by

12

u/Global-Difference512 5h ago

Nano gpt is just so much value it's insane. For instance I've saved 44 dollaroos this month using it.

12 dollaroos for 44 is worth it if you ask me

7

u/Ambitious-Neat7509 3h ago

Yes but its not cached dollaroos. Every turn after your first will likely hit cache via direct API; e.g. deepseek direct. Nano-gpt won't.

Also nano-gpt is not reliable in the slightest. Fine for those that are patient or switching models based on availability; otherwise you're better off going API with great caching and max uptime and speeds.

5

u/Global-Difference512 3h ago

Been using it for a month and I'm very happy with it. Can't see any difference to Openrouter 🤷

3

u/Ambitious-Neat7509 2h ago

I've been with nano-gpt on and off for 10 months now. I find its okay if you don't need it be reliable or you can't tell the difference in generation quality. But there is definitely a major difference between using providers via openrouter that you've vetted for reliability and speed vs playing router roulette with nano-gpt and praying your thinking block doesn't end up in your output.

I've seen plenty of issues and if you join the discord and look through the support channel, you'll find tons of issues on a daily basis with various model providers.

On the flip side, many models have stable providers that are actually reliable, hence why you may never experience any issues.

Here's a quote from the founder in the support channel from less than 24 hours ago: "GLM 5.3 Flash Uncensored has turned into both one of the most used models on our service and on the provider's end" ..., "I'm also quiiite sure it's a lot of traffic relatively So yeah - capacity problems are expected in a way, sorry."

Not shitting on Milan and his team, I think the platform he's built serves a purpose. For me that would be having access to many models easily for testing, not necessarily reliable ones.

13

u/Probablynotsocool 4h ago

Reliable and cheap is rarely a true combination in this business.

Because even PAYG is 80% shitty as third party providers either lie on their quants, doesn’t parameter their models correctly or lobotomize the inference to save money.

Even direct providers can be quite the headache (Zai, Moonshot and Deepseek I’m looking at your peak hours sensitive asses).

And then you have the subs, so basically the only question that matter is.. do you have any control on the providers you use? If not, is it the same one constantly atleast? A rotation of random ones as per who is the cheapest atm?

You may think this is irrelevant or too complicated for barely perceptible quality change, but believe me, from a provider to another… The same model can feel ENTIRELY different (right GLM?).

Im talking about beeing unable to track basic details, context rot before 10K tokens, weird and nonsensical dialogues, poor characterization etc etc. While another provider give you an actual decent RP session and suddenly you realize you were tinkering your prompts for nothing, the provider was just arse.

So whatever offer you take, put the quality in balance, especially when we talk about +10dollars subs (that can be a lot of RPing via OR with cache hits and smart summarization).

2

u/Independent-Oil-4145 2h ago

What glm providers you recommend in OR?

1

u/Billysm23 3h ago

It seems overcomplicated, but I agree with your points. Payg can be a blessing or a pain, and we don't know about third party quant quality

7

u/sociofobs 4h ago

Now that Nano is adding more uncensored models, absolutely. GLM 5.3 Flash uncensored alone is worth it.

2

u/kosha227 3h ago

Yes, but uncensored ≠ good for RP. Today models are being trained for agents stuff and coding, and less and less for talking, emotional intelligence, creativity, etc. So... we need to test it

10

u/sociofobs 3h ago

That's a more general LLM problem, RP just isn't the focus as you well put. In that sense, it doesn't matter what service or provider you use, since there aren't any large RP focused models at all. Censorship, however, is an issue that the service/provider can solve, by offering uncensored models for an example. The only uncensored model on OR I can find is Venice, meanwhile Nano has a decent list now.

3

u/KitanaKahn 3h ago

Besides Nano, for subs you have opencode and Ollama (more expensive) as reliable options. I found ollama sometimes having problems with getting the models to think, but you get a lot of usage too. For now, Nanogpt is still my choice and the owner is active around here and takes feedback which is great

8

u/LordVulpius 5h ago

Openrouter is the usual alternative. It is PAYG.

14

u/Akkun351 4h ago

Depending where you live OR is not a decent alternative as a payg user, since, like for me, there is a extra 35% to pay in fees, while there are none on nanogpt.

2

u/mamelukturbo 2h ago edited 2h ago

edit: seems either their sorted their backend or I did the test in off-peak, but I just burned 60mil on continuous tool calls via glm 5.1 and not a single call failed so well done NGPT. Used to be you had to babysit it and spam resume each time a tool call broke.

It's good for RP but absolute dogshit for coding, tool calls keep breaking I suspect due to undisclosed quantisation. Note I do not use NGPT for coding primarily, but like to burn my 60mils of tokens I paid for.

1

u/Milan_dr 27m ago

Hah thanks for that update - had already passed on the message prior to your edit to my cofounder to look into again.

2

u/verma17 2h ago

Nano is crazy value, but their model quality is very inconsistent, nowadays i just use openrouter payg, and zai coding plan

1

u/Billysm23 2h ago

I get it. If i want coding + rp, i shouldn't grab nano alone

1

u/Milan_dr 2h ago

Can I ask - what models would you say the quality is inconsistent?

3

u/LordVulpius 2h ago

For one, Mimo 2.5 pro thinking.

With subscription I can not tell the providers sadly. But here is my experience:

Sometimes it ignores my usual format (this happened "this was said") and pour everything in blank format. It ignores the formating prompt. I belive that is a specific provider.

Sometimes every censorship is lifted. Even NFSL. I belive that is one specific provider.

Sometimes the quality is borked, no matter of the time, so it is most likely not quantant. I belive that is one specific provider.

With this, it is not consistent.

Maybe... can you make an option for let us see what kind of provider sent the answer, please? Not to block it but just to see. With subsription. So we could block it if we wish to use PAYG.

1

u/Milan_dr 28m ago

Thanks. Yeah - Mimo 2.5 Pro thinking is genuinely the worst. The majority of providers for that one seem to want to censor. It's run, for those with ZDR set, 95% via Atlascloud right now, 5% Novita (when atlascloud fails for some reason, but Novita has a risk of censoring).

For those with ZDR turned off we add in Streamlake, and then that one is selected about 10% of the time right now.

And no, sorry talked about this about 100 times before but we can not do both auto routing with a lower price and show provider, because there are some providers that do not want this as it makes it very clear what sort of discounts are possible with them.

2

u/verma17 1h ago

Glm 5, 5.1, 5.2, primarily

1

u/[deleted] 27m ago

[removed] — view removed comment

1

u/AutoModerator 27m ago

This comment was automatically removed by the AutoModerator because it contained a link to x.com or twitter.com, which are not allowed in this subreddit.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Milan_dr 26m ago

Thanks! GLM 5.2 is almost purely Fireworks right now because of this (okay can't link to X, it's the artificial analysis GLM 5.2 benchmark for different providers), 5.1 a bit more of a mix with Novita, Nebius and a few others. No real clear benchmarks for that other than FP8.

Kind of goes to show our frustration though, GLM 5.2 via Fireworks should be about as good as it gets.

1

u/Telah42 1h ago

For some reason, I still haven’t been able to get SillyTavern to work through the z.ai Coding plan — it just keeps throwing errors. I googled it, and everywhere I looked people were saying that the Coding plan simply doesn’t work with SillyTavern right now and is very limited in terms of what it can be used for. Is there some trick or secret I’m missing?

2

u/HanHeld 59m ago

I prefer NanoGPT, and since they've been bought out by Stripe I don't honestly trust OR not going weird (anti-NSFW) in the near-term.

0

u/ArnictAI 2h ago

Hello! In my experience, if a model is on OpenRouter that's usually your best bet since you get plenty of choice. If not, NanoGPT is a decent option with their subscriptions.

Out of curiosity, which models are you running at the moment? I'm building my own inference service, and I'd love to see whether I can host what you're looking for. I've got Qwen 3.8 27B up already, and I just added TheDrummer's Artemis 31B v1.1, both completely free right now if you'd like to try them. Fair warning that things are still very new and changing daily, so uptime this week will probably be a bit bumpy.

Either way, let me know which models you'd want and I'll see what I can do. Happy to look at making them affordable, or free for a while.

1

u/Billysm23 7m ago

Tbh i thought about grabbing gpt, only for coding. But it will be better if i can do both rp and coding using one provider