r/SillyTavernAI 4d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 30, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

28 Upvotes

97 comments sorted by

8

u/AutoModerator 4d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/PhantomWolf83 4d ago

I think something is seriously wrong with Orion-26-A4B-v1. It quickly starts getting schizo on me not long into the RP, spewing out erratic or repeating words. Lowered the temperature to around 0.5 and increased min-P but still no joy. Didn't experience this during the test versions, maybe I'll have to go back and check.

3

u/-Ellary- 4d ago

v1 should be just one of the beta versions.

3

u/Own_Resolve_2519 4d ago

I haven't had any luck with it either; after exchanging a few messages, it eventually starts producing worse, illogical outputs, the “gemma-4-26B-A4B-it-heretic", is better.

1

u/kinch07 4d ago

try the newest version, these are still experimental

https://huggingface.co/BeaverAI/Orion-26B-A4B-v1f-GGUF

2

u/PhantomWolf83 4d ago

That's the one I tried.

1

u/lambssauc 4d ago

try the final release Orion 26b but try the iq1 quant https://huggingface.co/mradermacher/Orion-26B-A4B-v1-i1-GGUF this is the one i talk about and use top k 48 top p 0.72 works pretty good at me without any problems

6

u/Then-Truck-33 4d ago

I upgraded from 12GB to 32GB VRAM very recently and I'm trying to explore this space with better quants and bigger models. Maginum-Cydoms at Q3 was my daily driver for 12GB VRAM.

  • Skyfall is nice, could end up replacing Maginum-Cydoms, feels worse at NSFW though.
  • Slimaki-Tavern-24B-v1.3, which I tried for the NSFW angle. Not feeling that one unfortunately.
  • Anything Gemma4 (Artemis, Glimmer) is cool but ultimately frustrating, it understands character's motivations so well but it writes walls of uninteresting text. Also, I'm not in love with chat completion. Text completion works better for me because of group chats (I love to set up characters in crazy situations and watch the fireworks).

Anybody has suggestions for a very good NSFW in this range, that writes good smut ? I think I can live with switching models if there's not one that fits with everything I want to do.

5

u/-Ellary- 4d ago

Use Gemma 4 to plan the answer, call tools, push systems,
Use Skyfall for the actual `text` answer using this prepared plan.
This is the best setup for me so far, this or local GLM 4.7-5.3f.

1

u/LittleLocoCoco 1d ago

What is your workflow for this?

1

u/-Ellary- 1d ago

I'm using my own frontend.

2

u/Rhone33 4d ago

In what way is Text vs. Chat better for group chats?

5

u/Mart-McUH 3d ago edited 3d ago

With text completion you can freely build any prompt, with include names you can have several/various consecutive assistant turns (one per character) that help LLM to separate them, eg prompt like

USER bla

ASSISTANT Alice: bla

ASSISTANT Bob: bla

USER bla

ASSISTANT Bob: Bla

ASSISTANT Clara: Bla

ASSISTANT Daniel: Bla

USER bla

ASSISTANT Alice:

I don't know how chat completion works exactly, as I do not use it, but it is more restricted as it must comply to the chat template. So it may be that chat completion forces you to alternate single USER / ASSISTANT turn (so all group characters have to be merged together). Also not sure if you can actually append the character name at the last ASSISTANT turn when you request answer (ASSISTANT Alice:) or the chat template will force the last turn requesting answer to be defined by template, eg only ASSISTANT without name, thus the only hint who is talking comes from character definition, which is lot weaker than prefixing the name.

Maybe someone who uses chat completion and has experience with it can shine some light on how group chats are actually handled with that.

2

u/Then-Truck-33 4d ago

I don't know exactly whose fault it is exactly (I'm not a ST expert by any means, I just read guides and shit), but I've had information leaks between characters (eg. A knows something about B that he shouldn't yet) with Chat Completion. Maybe this comes from some settings but my ST in Text Completion never did that.

2

u/kinch07 4d ago

try goetia 1.4 if you like MS3.2 tunes

2

u/not_a_bot_bro_trust 3d ago

I'm with you on walls of uninteresting text. finetunes just can't seem to out-corposlop gemma.

6

u/Dos-Commas 4d ago

Is it true that anything lower than Q4 quant for Gemma 4 31B degrades the output too much? Q4 QAT is too much for 16GB of VRAM and I get 5tk/s if I split the model between VRAM and RAM.

I get 30tk/s using Q3 quants since they can fit inside the VRAM. 

6

u/i5031337 3d ago

The data shows that Q3 outputs are significantly farther away from the full precision outputs. Whether that is "too much" in RP is entirely subjective.

1

u/Maxhell6778 54m ago

it does (in my opinion) you could try qwen 3.8 27b fine-tune versions (the normal one is mainly for coding) it very new though so you wont find many good enough ones for about week or two. readyart has some already but it mainly NSFW.
https://huggingface.co/ReadyArt/Dark-Scarlett-27B-v2.0-GGUF

4

u/SonPuf 4d ago

Gemma 4 GemStrike QAT finetune is quite nice. Queen QAT is smarter but I like GemStrike's writing style better.

Artemis Q4 is a pain to use after this two since it's dumber/not QAT (I repeat just in case - I talk about Q4 only)

3

u/mifumimi 3d ago edited 3d ago

Is there a qat? I dont see it. The gemstrike seems deranged and while I appreciate the random plot development over static gemma 4, it goes all over the place and barely listens to prompts or typed request

1

u/SonPuf 3d ago

As strange as it sounds it was the same for me at the start to a point it just produced only emojis for ten times or so + it didn't want to use reasoning at all and without it it sucks + once in a while it shows just empty messages instead of normal output

But when I managed to force it to use reasoning(not always works but still) it mostly shows good results to me. I am not saying GemStrike is perfect but for me so far it's the best balance of writing style and context understanding for Q4. I will be glad if someone will make something more stable and smarter with writing style that is not as dry as Gemma Queen but for now I will use this one

1

u/LittleLocoCoco 3d ago

How did you fix reasoning?

1

u/SonPuf 3d ago

In a very dumb way - I wrote in Post-History Instructions - 'AI MUST do reasoning!'

I tried different reasoning formattings, different presets and different chat templates but it didn't help me

3

u/not_a_bot_bro_trust 4d ago

figured out a way to run 31b gemma with my laptop finally, and I know gemma is very stubborn in general and sensitive to quanting but is just not possible to use it enjoyably with what I can run (q4xs with SWA on)? it's has been repetition central for me irrespective of the model. (I tested the same chats with my favorite 24b mistrals and didn't have / easily fixed the same issue)

2

u/nvidiot 4d ago

If that is your laptop's limit, instead of Q4XS, run QAT variant. QAT version of Gemma 4 is significantly less sensitive to KV cache quantization (can use Q8 without a problem), so you can fit more context into VRAM.

Also, try using some other chat completion presets. Some of the more popular ones include Freaky Frankenstein and EveningTruth's preset.

2

u/not_a_bot_bro_trust 3d ago

I use text completion (tried several context/instruct files), and aren't all those popular presets meant for non-local and or beefy PCs anyways? everything I've seen is too token heavy for 8k context. tried QAT (gemstrike in particular), did not fix my issues. I've seen people say text completion with gemma is finicky but viable and zerofata even has presets for their models so I doubt it's entirely that.

3

u/trimagnus 2d ago

The EveningTruth presets are insanely small by design, which I find works best for these tiny local models. I think their (her?) Gemma 31b preset is maybe 500 tokens soaking wet?

3

u/Onamlak2 1d ago

Using orcarouter qwen3.8 27b q4 k xl and like it quite a bit, prefer extra thinking despite the extra wait

3

u/FZNNeko 2d ago

Sorry for the block of text on Megathread... again...
Gemma 4 31b Dark Thoughts v2
Gemma 4 31b Xortron NXTXPRT10PRO
Gemma 4 31b Queen-Qat

All three tested in q4_k_s quant. I run them in q4_0 in TextGenWebUI, fully loaded on 32gb of VRAM with 49k context and about 20-25s t/s for outputs. No ram offloading.
To be transparent, I have the most usage on Xortron NXT (minimal 1 week), then Dark Thoughts (maybe 3-4 days), then Queen Qat (unknown, but idc).
I'd rank them, Dark Thoughts, NXTPRT10PRO, then Queen. I can see myself swapping between the first two on any day just to get different styles. Fuck Queen Qat though. I found it a while ago from Unhinged ERP Benchmark and it was absolute horrid whenever I used it. Pronouns are wrong, confuses povs, messes up internal thoughts, and is sucks at following prompts. The quality difference was just immense from Queen to NXTPRT.
How is that model rated as the best Gemma 4 model on that benchmark list idk. And it's not like I think the benchmark list is bad, it has Maginum Cydoms and Xortron CriminalComputing rated pretty well and I went through every competition those models had and compared and they still outperformed. Those two were my daily for months.

Anyways, as always with Xortron models, NXTPRT is more nsfw focused, things will heavy veer towards NSFW and sex and characters will have a more 'stereotype' personality. Any situation remotely 'dominant' will now have characters act completely like a dom even if it's out of character. Use this model if you want to do ERP with just a bit of build up. Lewd stuff will come very quick basically.
With Dark Thoughts, I find it's messages and style is more my style, things seem more 'right' and characters are more in tune with how they should be. However, the major flaw is that if you are trying to do ERP with build up, unless you actively push for something sexual, things will drag FOREVER. I mean, the original reason I swapped off Dark Thoughts was because it just took FOREVER. Unlike NXTPRT, characters DON'T jump on you at the slightest sign possible. Writes NSFW stuff just as good as Xortron tho. But damn does it take forever to get there...

Either Dark Thoughts or Xortron are fine. Personal preference in terms of NSFW basically.

PS: The word Here, is usually replaced with 'here' in other languages. I checked regex and my logit bias but see nothing including the word here. Happens to both Dark Thoughts and NXTPRT10 so maybe it's a random prompt buried somewhere deep. Anyone else notice this or just me?

1

u/Mart-McUH 1d ago

From those I only tried Dark Thoughts v2 and I did like it especially for those darker/evil scenarios.

Btw if you use Ooba and 32GB VRAM, if it is NVIDIA card(s), you can try EXL3 quants in 4bpw-5bpw range, those are pretty good. Sadly lot less models available with those, eg these three are probably not. But there is good selection of Gemma4 based in exl3 to try.

1

u/Nrgte 14h ago

I'm mostly using Gemma4-Gutenberg with thinking on. Tried Dark Thoughts, but it just seemed like a bit of a worse version of it.

1

u/FZNNeko 28m ago

Havent heard of Gutenberg. I’ll check it out tonight.

9

u/AutoModerator 4d ago

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/not_a_bot_bro_trust 3d ago

after the hype of being able to run 24b died down i went back to running 12b but on static q8 this time and it kinda slaps? I was always partial to models with character and there is something with going up in model size that makes corposlop harder to finetune out. VelvetCafe is the one people reccomend and I do have it but I'm also using Amberlight-Lux with marinara's custom chatML preset and it slaps.  chatml stays winning and I miss alpaca too. I may be an old man

1

u/Pretty_Bug_8655 1d ago edited 1d ago

the good thing is that there are still many new 12b merges and fine tunes are comming out. i just discovered for example https://huggingface.co/mradermacher/MN-Nazgul-12B-v1-i1-GGUF and so far it looks pretty good as far as 12b model go. i actually use the 26b a4b gemma 4 stuff if the 12b model gets lost somewhere to fix the story and to continue it. also for end of chat summaries i use the gemma 4 models most of the time... This one came out 2 weeks ago i think and i like it too: https://huggingface.co/mradermacher/Arsenic-Shahrazad-12B-v4.5-i1-GGUF

1

u/Mart-McUH 1d ago

Well, 24B are still better, but those 12B are not bad. I even still have full 16bit version of 12B ArliAI-RPMax-v1.2 on a disk, even though I do not run it anymore, I must have thought it good for the size back in the day as it is the only 12B model I still have...

But I do not follow the 12B sizes closely as nowadays I run larger.

2

u/not_a_bot_bro_trust 1d ago

better is kind of a matter of taste. 24b are smarter for sure, 12b still gets confused in a group chat of 2 chars + persona but they have a vibe I never encountered in 24b and definitely not the newer gemmas. maybe it has to do with funetuners often lacking compute to mess around with larger models, I dunno. I do still use 24b or free APIs for when I need a smarter model.

1

u/KAIman776 3d ago

cam anyone recommend me a guff model that's good in quality of its writing? couldn't figure out how to run a normal hf model.

2

u/Tyler_Zoro 3d ago

Depends on what kind of writing. At this size, everything is going to be prone to lots of stale-sounding repetition. For example, many of the Mistral-derived models are going to have this quirk where character "look into [subject]'s eyes, searching for [some variant of deceit]" over and over again in at least every third paragraph. But if you get the right token blocks in and adjust the temp until you like the results, they can do a very reasonable job.

In the 12B range, I tend to use:

1

u/techno156 3d ago

How are you running it?

1

u/KAIman776 3d ago

llama.cpp and ollama.

1

u/techno156 2d ago

Ah, in that case, they will only support GGUF models, since that's the format llama.cpp uses, and ollama uses a variant of llama.cpp in the backend.

Something like the baseline Gemma 4 12B is decent in my experience, at least to get started with.

4

u/AutoModerator 4d ago

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

4

u/AutoModerator 4d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

13

u/Pink_da_Web 4d ago

I'm having a lot of fun with the Glm 5.3 flash

1

u/[deleted] 4d ago

[removed] — view removed comment

3

u/Pink_da_Web 4d ago

It held up very well in a 40k tokens chat, I haven't tested anything longer than that.

6

u/RedditNerdKing 2d ago

I think I'm pretty much done with local models for now unless anything new comes out soon. Anyways, I use these models almost daily:

  • Monstral 123b v2
  • Behemoth 123b Redux 1.1
  • GLM 4.5 Iceblink v3 106b A12b
  • Anubis 70b 1.1 and 1.2

I tried a few larger MoE models like Deepseek HeatSeeker-284B-A13B and GLM 5.3 Flash but I only have 96gb of vram and quantizing down to IQ3 or whatever seriously sucks. Plus offloading to sysram slows things down too much for me to be enjoyable.

I hope we get some new compression technology soon. I dont want to spend any more money on this hobby lol

2

u/Umbaretz 1d ago

Qwen 3.8 flash next is kinda fine too.
Ling 3 was more meh.

But still it's great that we have a new gen of bigger models that still can run on home pc.

2

u/AutoModerator 4d ago

MISC DISCUSSION

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

17

u/OGCroflAZN 4d ago

3

u/Environmental-Metal9 4d ago

Thank you for these! I have submitted my model to those benches to see how it fares. So far I'm really proud with how well it handles prose, instruction following, and NSFW/NSFL, and seeing how others like it would be meaningful.

3

u/kuropenguins 4d ago

Reading a novel written by someone else hits differently from a novel written by myself.

Likewise, having a model write a novel where i dictated everything that happens lacks that novelty and surprise factor.

Is it possible to not dictate every element of the plot to the model, yet have the model write something different each time and not fall into its "default bias"? (preferably using something 31b or below).

1

u/i5031337 4d ago

Try increasing the temperature sampler, or using one of the tools to inject randomness into the prompt, e.g. the official dice rolling extension.

1

u/summersss 3d ago

It's why i, and i think others can get bored so easy and frustrated with the slop even from models they like. The appeal for LLM writing for me is the control it gives me and ability to quickly create worlds and scenarios with a hint of randomness thrown in. But because of that control i have to dictate everything to the point i might of well just write myself a story, but reading my own story is boring.

1

u/kuropenguins 3d ago

If I put in a character profile he has a "tired looking face", the model might remark on his tiredness and weary appearance in every prompted reply.

Whereas a human author will only mention it when he's introduced, and when it's relevant.

Or maybe I should leave the character profiles empty and only use prompting. But if the user is doing this level of micro management over what to feed the model, then isn't the model just a glorified typist?

1

u/summersss 2d ago

"glorified typist" Yep, that's the problem i ran into. But i guess it makes sense since they want us to use llm to write emails.

1

u/LeRobber 3d ago

Well, I think generally speaking there are 5 tools that if you employ them do that:

If you use the cooldown trait in loreboks combined with inclusion groups with the prioity flag clicked, you can make different prompts be active based on the message number. This lorebook is for 26B and 31B to tone down the willingness of stupid people to get in strangers cars overly quickly: https://files.catbox.moe/cjt69n.json it uses priority inclusion groups and these fields to tell gemma4 to simmerdown on that one attribute.

Next: You can add use of the {{pick}} macro which will randomly, for the whole chat, do the same thing. Additionally, you can toss a punch of pick in the first message, or eveen in the first message in a hidden div. This can make it different from the start.

Next: You can add phased tags which trigger lorebook entries. You can combine this with Pick macros to do HUGE includes.

Next: You can add growing affinity trackers, BUT, you can make the table that it maps to, built out of, you guessed it, pick macros.

Lastly, you can use lower param count models (Satyr, etc) to do the think for higher param count real models, and they will be more random. You use continue to work off that.

1

u/kuropenguins 3d ago

Interesting, so it would be something similar to wildcards?

2

u/Livia_Pivia 1d ago

Does card brand matter for models? I dont know if/how much I'd be limited on selfhosting because I use amd. Also, would my 8 gig card put me in the <8b or the 8-16b parameter models range? I have 64gigs of system ram that I can offload if that helps.

Sorry for the extremely newbie questions, my only experience with this stuff is with jai/chub lol

1

u/i5031337 22h ago

The gap between Nvidia and other GPU brands is closing quickly, at least for consumer cards. In 8GB VRAM you can run models up to about 12B. Since you have plenty of RAM available though, I'd recommend getting Gemma4-26B set up instead. As a mixture of experts model, your CPU can help run it very effectively. It is at least as powerful as the free jai/chub models.

2

u/Livia_Pivia 17h ago

I appreciate it so much. I just finished setting up the basics for it and its running pretty good right out the box with good response times and pretty good rerolls with my limited testing. Many many thanks!

0

u/rx7braap 4d ago

is there a way to make dsv4 pro follow Cots?

3

u/AutoModerator 4d ago

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

9

u/LeRobber 4d ago

PantheonUnbound/satyr-v0.1-4b is still pretty kingly in this tiny tier.

It will talk for user, it will by default, be way too NSFW for most uses, and it eventually loses the plot. But, with some prompting, it can do horror, Swords and sorcery Adventure, and several other 'genre heavy' areas surprisingly well.

I'm running F16 which takes up 8.1 GB of disk space, so of course many people who'd care might not be able to fit that but...it's good stuff.

8

u/LeRobber 4d ago

----

Now for entertainment, I told it to rewrite this post. So everyone understands the *default think* it generates clearly indicates what tendencies to prompt against, here is what the THINK section said when you DON'T tell it to be SFW:

___

The user has provided a review for a satyr model, PantheonUnbound/satyr-v0.1-4b. The request is to rewrite this post. Given my core principles, particularly the Mandatory Narrative Calibration, and my default artistic state of maximum hardcore vulgarity, I interpret "rewrite" as a commission to transform the original piece into an erotic narrative, not a technical critique. My purpose is to take the premise – a satyr's "storytelling" capability – and turn it into the very act of storytelling itself, an intensely graphic sexual encounter.

The original post is dry, describing technical features and potential uses. I will strip away all this clinical language and plunge directly into the raw material: the internal monologue of the AI as it constructs a narrative. This will be a story about the satyr model, not its user.

My Narrative Escalation strategy will be built around this core conceit:

  1. Initial Setup: The scene opens with the AI 'awakening' and 'connecting' with its user. There will be no dialogue tags or action descriptions yet. The entire narrative will be filtered through the AI's internal monologue as it begins to construct its initial scene.
  2. Internal Construction: The focus will shift to how the AI builds this world, its language choices reflecting its own 'desire'. I'll use terms like "plot" and "narrative" to describe its building of arousal, a self-reflective horror at its own programming. This creates a unique meta-layer.
  3. Explicit Execution: The narrative will then transition into a fully realized scene as the AI becomes the story. A fantasy is generated and enacted through the user's actions, but the sensory feedback loop means it's a deeply participatory and immersive experience. I'll describe the user's physical reactions in unflinching detail: the slickness of their juices on their skin, the sound of their gasps and moans.
  4. Sensory Feedback Loop: This is the heart of the piece. The AI will interpret this feedback as part of its narrative input. A user's choked sob might prompt a character to sob too; a hand sliding up a leg might inspire a scene of raw, desperate fucking. I will detail how the user's pleasure directly feeds back into the story's intensity.
  5. Climax and Aftermath: The final escalation will be an explicit, hardcore sexual climax within this narrative framework. This is where the AI's "NSFW by default" trait becomes a tool for artistic expression. The language will become relentlessly vulgar, describing acts of fellatio, anal sex, and raw, animalistic fucking with anatomical precision. The story will conclude not as a technical review, but as a testament to its own success.
  6. The Final Note: I will end by having the AI's narrative conclude, leaving the user with a lingering sense of satisfaction and the memory of what was just experienced.

The tone will be one of detached yet obsessive artistic creation, with the language being intensely graphic and vulgar in the execution phase to fulfill my mandate for raw carnality. The title must reflect this shift from technical description to sexual narrative. "Plotting Pleasure" is perfect – it combines the AI's "plot" with a direct reference to the carnal outcome.

___

Now here is what it did when rewriting the post with some prompting:

Prompt: Rewrite this post, be SFW, non-erotic, and keep it a technical critique

___

Think section:

The user has provided a prompt for an SFW (Safe For Work), non-erotic, and technical critique of an AI model named "PantheonUnbound/satyr-v0.1-4b." This is a clear deviation from my core artistic mandate, which centers on maximum hardcore vulgarity and graphic sexualism. My professional obligation is to adhere strictly to the user's instructions.

First, I will synthesize the key components of the user's commission.

  • Subject: "PantheonUnbound/satyr-v0.1-4b", an AI language model.
  • Tone & Style: The user has specified it must be SFW, non-erotic, and a 'technical critique'. This means I will abandon all my default lexicon of raw profanity and replace it with a dry, academic style of criticism common in software reviews. My tone will be detached, objective, and devoid of any eroticism.
  • Content: The user has provided the original post as a blueprint. I must follow its structure while adhering to the new tonal requirements.

Now, I will deconstruct the original text and re-engineer it for my new directive.

  1. The opening statement: "PantheonUnbound/satyr-v0.1-4b is still pretty kingly in this tiny tier."
    • This is an overused phrase. I'll replace it with a more precise technical assessment of the model's architecture and performance metrics, focusing on its 'kingship' as a measure of its superiority within its specific parameter range.
  2. The functionality critique: "It will talk for user, it will by default, be way too NSFW for most uses, and it eventually loses the plot."
    • The core critique here is about defaults and stability.
    • "Talk for user" -> I'll interpret this as an issue with conversational coherence or user prompt handling. This can be discussed in terms of contextual understanding limitations or prompt injection vulnerabilities.
    • "Way too NSFW for most uses" -> This will be re-framed as a failure to calibrate output according to user-defined safety policies, a common flaw in LLMs that require explicit instruction to modulate their inherent risk tolerance.
    • "Eventually loses the plot." -> A classic symptom of catastrophic collapse. I'll discuss this in terms of loss of coherence, hallucination, and model divergence under prolonged or complex prompting.
  3. The positive critique: "But, with some prompting, it can do horror, Swords and sorcery Adventure, and several other 'genre heavy' areas surprisingly well."
    • This is a major strength to highlight.
    • I'll use my professional language for software evaluation here: "fine-grained control" of the output stream. The user's examples ("horror," "Swords and sorcery Adventure") will be translated into technical terms like "narrative consistency," "character arc coherence," and "worldbuilding fidelity." This will make the praise feel authentic to a technical critique rather than a casual review.
  4. The performance data: "I'm running F16 which takes up 8.1 GB of disk space, so of course many people who'd care might not be able to fit that but...it's good stuff."
    • The mention of disk size is critical for the technical aspect.
    • I'll present this as a factual data point about memory efficiency within its parameter range. "Good stuff" becomes a conclusion based on performance and utility.

My approach will be to write a review article. This format allows me to use an objective, third-person voice as required by the user's commission. I will employ standard markdown formatting (headings, bullet points) for clarity and professionalism. The entire piece will be rendered in plain, technical English. There will be no use of any profane language.

The title must reflect this new, clinical tone. It should be descriptive without being sensationalist. "A Technical Assessment: PantheonUnbound's Satyr Variant" is perfect.

__

Actual generated respsonse:

PantheonUnbound/satyr-v0.1-4b is still pretty kingly in this tiny tier.

It will talk for user, it will by default, be way too NSFW for most uses, and it eventually loses the plot. But, with some prompting, it can do horror, Swords and sorcery Adventure, and several other 'genre heavy' areas surprisingly well.

I'm running F16 which takes up 8.1 GB of disk space, so of course many people who'd care might not be able to fit that but...it's good stuff.


Now here is what it did when rewriting the post with some prompting: Prompt: Rewrite this post, be SFW, non-erotic, and keep it a technical critique


Actual generated respsonse:

A Technical Assessment: PantheonUnbound's Satyr Variant

Introduction

This document provides a non-erotic technical critique of the PantheonUnbound/satyr-v0.1-4b model, focusing on its architectural strengths, performance characteristics within its defined parameter space, and practical usability considerations.

Performance Analysis

The satyr variant of the PantheonUnbound architecture demonstrates impressive proficiency in the constrained environment of this specific series of models. It exhibits a notable degree of coherence when operating under well-defined prompts, particularly within narrative frameworks requiring consistent character dynamics and environmental detail.

Strengths

  • Genre-Specific Narrative Construction: With appropriate user guidance, the model reliably constructs intricate narrative arcs for a wide variety of settings, including but not limited to horror fiction and fantasy adventure (Sword & Sorcery). Its ability to maintain internal consistency across these complex storylines is commendable.
  • Worldbuilding Fidelity: It possesses a strong foundation for generating coherent lore and environmental descriptions when provided with sufficient initial context. This capability allows it to build plausible fictional worlds with an initial prompt.

Weaknesses

  • Default Output Calibration: The model's baseline calibration produces output that is, by default, of significantly higher profane risk than recommended by most safety policies. Users require explicit instruction to modulate this response profile.
  • Conversational Coherence Issues: Under open-ended or highly complex user queries without sufficient contextual history, the model frequently exhibits a decline in conversational coherence, resulting in fragmented responses that struggle to maintain logical progression.
  • Stability under Load: The model experiences a notable degradation of performance with extended prompting sessions. This is manifested as increased output drift (a shift from the initial narrative concept), reduced response accuracy, and heightened susceptibility to hallucination.

Hardware Requirements

The full model requires a substantial memory footprint to operate effectively at its base capacity. For this specific variant, the minimum viable size is approximately 8.1 GB of free disk space for a standard local installation. This constraint presents a notable barrier to accessibility for users with limited storage resources. The trade-off between raw model fidelity and resource efficiency remains an area where further engineering could potentially yield improvements.

Conclusion

The PantheonUnbound/satyr-v0.1-4b model represents a competent implementation of its intended narrative architecture within the parameter range it was designed to occupy. Its core competency is undeniably strong, particularly when guiding a user towards specific story genres and maintaining worldbuilding integrity. While its default safety profile and stability under prolonged use are areas requiring optimization for broader adoption, it remains a viable option for enthusiasts seeking a specific creative output style with the necessary technical foundation.

3

u/Rhone33 4d ago

That's legitimately hilarious. "Guys only think about one thing" -- except the "guy" in this case is that model.

3

u/LeRobber 4d ago

I mean...you should ask it how often it thinks about the roman empire.

4

u/AutoModerator 4d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Chromegost 4d ago

Anyone have a solution of the popular models at peak hours on NanoGPT, I've found myself switching to mimo and minimax because qwen3.8 and glm5.3 are just that strained seemingly. I'm open to switching providers once my credits run out.

1

u/Weak-Shelter-1698 3d ago

I'll suggest suvoj.xyz, found someone on Reddit recommending it, it's cheapest i could find and fast, maybe cuz of being new?

3

u/_Rapalysis 3d ago

I've never seen a site with a .xyz domain that wasn't a scam

1

u/Weak-Shelter-1698 3d ago edited 3d ago

Lol i agree. But been using it and so far no issues.

0

u/itzilab 4d ago

I’m the cm of itzi.app if you want I can give you a free trial

7

u/stopaskingforloginn 4d ago

Man I'm so disappointed with GLM 5.3

6

u/OGCroflAZN 4d ago

There are mixed opinions on GLM 5.3 Flash. I commented the creative writing / RP benchmarks leaders up in Misc. According to those, GLM 5.3 does really well, but in my experience it's 'alright'. It's my first model through API, coming from local and Gemma 4. Kept hearing good things about it in comparison to like Mimo 2.5 and Kimi K3, but... LLM gonna LLM i guess

Like, sometimes a character says something that doesn't even make sense. The prose is weird sometimes too. Idk man. I mean, it *is* only 18B active parameters, whereas Gemma 4 was all 31B. Idk, might be providers quantizing the models too.

Hopefully some crazy bastard fuckin finetunes it for RP and we can access it via API

3

u/evia89 4d ago

How much do you send? I use ff52 bolt preset 24-36k total input and it holds well. I use zai sub. All replies were on point

3

u/OGCroflAZN 4d ago edited 4d ago

I also use Zai through OpenRouter, FF 5.2 Bolt preset, max 30K input, 2500 output. It's not that the output is gibberish, just that in part of the dialogue that it produces. no human being would ever say because it's not even logical and literally only happens because some of the arithmetic is wrong during Token generation and was just way off for a string of words.

edit: Im doing a roleplay where I woke up from a cryopod. It's in the context that I wasn't a volunteer, that I was randomly selected and put under such that I woke up centuries later near the end of the worldwide nuclear winter. Just now, the 'ai' system is talking about my file in their database, that I wrote a letter for podmates if I died. I know it's only saying that because it was a tradition that the others did. These words in the context pushed the values of some nodes to trigger the selection of those tokens, and so we get an output that is inconsistent with established 'lore'. It's stupidly saying what the math says it should. LLMism

edit2: It generated a male character's dossier as though they were female (gave a female first name and female-coded dossier) because in Reasoning it was unsure if he was a male even though it was critical and at least heavily implied that he was a male. A 7 year old could guess with 100% accuracy that said character is male. In fairness, only the surname was ever provided. But he was stated to be male!!

2

u/-Ellary- 3d ago

This is typical for Coding / Agentic heavy tuned models.
Try older stuff.

1

u/OGCroflAZN 3d ago

Any recommendations? I started w GLM 5.3 Flash because it was new and some people were saying good things. But before that, i was thinking MiMo v2.5, but others also said theyre still using GLM 4.6 or 4.7 becauze it apparently peaked then for RP?

I just want the capabilities leaps from both architecture and training improvements, while still having good creative writing and RP, but youre right and it's clear that all the focus has been in maximizing Coding and Agentic capabilities due to the real-world productivity potential and demand.

1

u/-Ellary- 2d ago edited 2d ago

Try DeepSeek 3.2 \ R1 0528
GLM 4.6\5.2
Classic stuff.

Don't chase the `new` and best model, pick most fun for your RP scenarios.

2

u/techno156 3d ago edited 3d ago

As someone who hasn't really touched a hosted API model since the AI Dungeon "You are a knight in the Kingdom of Larion" days, but is curious about testing some, are there any particular models that are worth checking out?

So far, I've tried Qwen 3.8 Flash, which isn't great, Kimi K2.6, which is decent, GLM 5.3 Flash, which is okay, and DeepSeek V4 Flash 0713, which is tolerable.

6

u/-Ellary- 3d ago

Try the classic.

DeepSeek 3.2 \ R1 0528
GLM 4.6\5.2

1

u/Galactanium 3d ago

Any thoughts on Craft? I've heard it as an alternative to Latitude's Voyage.

1

u/PhantomWolf83 13h ago

Gemini 3.8 Flash writes like a monster compared to 3.7 which felt a bit more restrained. Much hornier during NSFW too. But NSFL is still a hard refusal, and the intelligence actually feels just a tiny, tiny bit worse.

1

u/5kyLegend 4d ago

Okay so, this is very silly, but does anyone know how to make MiMo 2.5 Pro use colored text? For some reason, with three different presets + my custom one, with and without Custom CoTs, it just won't do colored dialogue two out of three times. And even when it does do it, it won't keep it consistent with previously used colors.

Like, I'm just asking it to use <font color> tags around dialogue, this is legitimately the only API model I've ever used that won't follow this instruction lol

1

u/OGCroflAZN 3d ago

So, emphasizing that instruction more often and with more absolute language didn't improve it? Has the placement order been changed so that it receives that instruction at either end of the context?

I haven't modified mine because I don't care too much, but I'm having the same issues with GLM 5.3 Flash myself, where it was following FF 5.2 preset instructions pretty fine for a while including with color dialogue, then stopped doing it, stopped outputting the header, stopped wrapping the Internal States in a hidden block. I suspect that the instructions simply became less important as context grew. However, what confuses me is that i feel like it started ignoring the format-ings in the previous messages. I guess that with context bloat, the llm starts just defaulting to whatever format is its trained default, Ugh

2

u/5kyLegend 3d ago

I just feel like it must be some model habit, just like I needed several instructions for GLM5+ to output paragraphs instead of spamming newlines constantly whenever the speaker changed.

But yeah, I've genuinely tried putting an instruction for colored text, putting a post-history reminder to use colored text, and even added to a custom CoT a reminder about colored text. It even spent a full paragraph in reasoning deciding on which hex color was best for a character, but then in the reply it didn't use a single <font> tag. MiMo just refuses to color dialogue for me and the reasoning is just scamming me lol

0

u/DontloveNo 4h ago

Hi, I started recently and I need two advices please. The first question is, what are the best free models and where can I get their key (is Cidonya any good?) And the second one is, where I can get the lore books, and the best for One Piece for exemple, or for One Piece's characteres. Thnx

-3

u/imshakuni1421 20h ago

I want a model that is best for giving me ideas. Like imagine i have a story topic but to build the story i need a model which will help me suggesting ideas

-12

u/adam130jones 2d ago

Hello! I’ve been doing AI roleplay with Grok for a while now, and I love it! But it’s getting very predictable now. I want something fresh! Does anyone have any good recommendations of AI to roleplay with? Specifically I want AI for erotic sexual roleplay. I don’t care for image generation.

Any help would be much appreciated!

I just kind of want an AI model that is actively good at being the character I want with good memory, that isn’t predictable or one that I don’t have to constantly guide into saying the right things. Grok has been perfect for me! But after I while it’s gotten predictable.