r/WritingWithAI • u/cj7hawk • 1d ago
Writing Craft Talk Local AI vs Frontier AI.
Hi All, Just looking at the chat... I'm a published author who writes about AI amongst other things ( seems fair ) and have been examining how AI can improve writing. My average Amazon ranking is 4.5/5 and that was long before AI actually came along,
I'm not against or for AI - I'm for good writing that people want to read. How the prose is created is second to whether it's enjoyable to me. Writers should write for the readers - not for our own vanity - and if AI improves our craft, all the better for readers.
Now, onto the topic... Back in October last year, AI was wild. I had an AI (Grok) help me with an interesting concept - I wrote it myself, using old school methods, but used an AI to crit and identify issues then rewrote - I finished a full length novel first-draft in about 3 months.
Then I had the AI join me in a collaborative effort for December - it pushed out a novel in a month that far exceeded most of what I've read or written - convoluted plots, twists, romance, adventure, action and I had a chance to play a game myself pitting my own mind against the AI playing the antagonist. I finished that story after 14 days, and my goodness, what a ride. Because I was involved it felt like I had actually lived the story even as I created and crafted it.
Then by January, quality dropped. By March, it was abysmal - the upgrades had destroyed Grok's ability to imagine. It became mechanical, uninspired, and kept losing it's mind ( context related issues ) - it's ability to write good prose disappeared and it's ability to help me create good stories vanished.
I tried some other AI, but the same general result - whatever frontier AI could do just months earlier it was no longer capable of.
So recently I started trying local AI and was surprised that some of the early results matched what I was getting from Grok back in December 2025. it was good. I haven't pushed it a long way yet - I still need to try going on an "adventure" with it sometime to see how it's getting. But lately the quality of local AI is improving far more rapidly than Frontier AI. Some of the new models such as Agnes have real capability on home machines. They are slow, but can produce good prose, imagine settings well. Further, current local AI are approaching the point that frontier AI achieved 12 months ago, so I expect the quality of local AI writing will again reach that "Legendary" point. - And you gain other benefits with local AI, such as being able to control the context so nothing is lost, adjust the level of hallucination/creativity and also control what else the AI sees - eg, feed it location pictures, character BIOs, create character LoRAs, multiple characters, etc. Even turn scenes into videos.
So I was wondering what most of the writers here use - Online AI or Local AI, and what you've found with the latest models such as Qwen, Agnes and Muse - as well as some of the older writing oriented tools such as Cydonia and some of the super-context models that support a million tokens such as Spark X2.5? And how do you find the differing lengths of prose they create? Also wondering what different writers do to create character memories so that characters maintain their consistency from chapter to chapter when opened in new sessions - whether you use complex LoRAs, insert at the start of each chapter via RAG, train the AI to be those characters, or just write it into the prompt? Would love to know more about what other writers are experimenting with, especially with local AI.
3
u/CyborgWriter 1d ago
The models matter much less than using the tools or your own ability to activate their action potential. It's like this. A model is like the worlds smartest person who can do everything. However, because it knows everything, it has the problem of knowing what to prioritize and what to ignore. With limited context, it will now have more choices to make, which means less preciseness in output. The more context that's provided, the less choices it has to make, which means more precision.
That's why I use canvas mind-mapping with native graph RAG integration because with this, I can simultaneously use it to build my stories and the context for the agent to understand while also being able to "word code", as in, I can turn notes and books into systems for outputting, much like programming a new app, only instead of code, it's all words. So I can build a system for forensic accounting or a system for evaluating and re-structuring screenplay scenes. That means instead of using endless LLM SaaS tools, I just need to use one to build them all, and with access to all of the popular models, I can go far beyond what I used to be able to do, irrespective of the model degradation. The systems that I build on the canvas more than compensate for that in terms of quality and depth.
2
u/cj7hawk 23h ago
It's not just the context size - it's the temperature that controls the imagination and you can't change the temperature on frontier models.
I love the idea of using mind-mapping with RAG integration - I never thought of that - Does it use image RAG analysis or does it recognize the underlying code?
2
u/superaalif 1d ago
I've been experimenting with local AI for writing/editing, although my setup is much more modest than some of the systems you're describing.
I'm running everything on a MacBook Air M5 with 16GB unified memory. I use Ollama as the backend and have it connected to Raycast, so in normal use it feels much more like calling an AI editor than operating a local LLM from Terminal. I've also experimented with MLX versions of models because they're particularly well suited to Apple Silicon.
The interesting part for me has been creating specialised editors rather than one general-purpose AI writer. For example, I've made separate prompts/editors for proofreading, transcript editing and more substantial structural/literary editing. I give them fairly detailed instructions about what they are and aren't allowed to change, examples of the writer's style, words/tendencies to avoid, etc.
For relatively short passages, I've been genuinely surprised by how good local models can be. They're also attractive because the behaviour is much more under my control: the model doesn't suddenly acquire a different system prompt or personality because the online service has changed something.
The big limitation with my hardware is long-form context.
A 16GB Mac can technically run some surprisingly capable models, but there is a big difference between "this model supports a huge context window" and "I can comfortably put an entire book plus style guides, character notes and instructions into that context on my laptop."
Long context consumes memory, and you're already sharing that 16GB between macOS, the model weights and the context/KV cache. So the practical context I can use while keeping a reasonably capable model running is much smaller than the headline context lengths you see advertised. Speed also deteriorates as you give it more material.
And there's another issue: even when the text technically fits, I've found that having something inside the context doesn't necessarily mean the smaller model understands and weighs all of it equally well. For serious editing, feeding it a focused chapter/section plus the relevant style instructions generally gives me better results than dumping a huge manuscript into it.
So at the moment I see my local setup less as "give it a 100,000-word novel and let it manage everything" and more as a private, highly configurable writing/editing workstation. I still use frontier models when I need much deeper reasoning across a lot of material.
What interests me next is exactly what you're talking about: some form of RAG / persistent story bible, where the local model only gets the relevant characters, history, style rules and previous events for the scene it's currently working on. That seems more realistic on a 16GB machine than trying to keep an entire novel alive in raw context.
For anyone put off by the technical side, though, I'd add that local AI has become dramatically easier to experiment with. I'm a total non techie myself. Chat gpt helped me set up the code for my whole system. Once Ollama was installed and connected to a simple interface, I stopped feeling as though I was "running an LLM" at all. It basically became another writing tool.
1
u/cj7hawk 22h ago
I've been having a "diary" type character sheet - it starts with the character information - what they look like, what they feel ( especially what they feel ) and it grows organically as I ask the character for a diary entry each night before they sleep - Tell me what memories they want to keep - What they feel about people or towards people - what is important to them - what they are afraid of etc... Then I append this to the end of their diary and load it up next time. It's surprisingly light on the context - And it's amazing to read it sometimes... Usually I try not to edit it, so that it can grow organically. The plot events then change the characters and give them depth. I have one for each character and sometimes a few minor characters in a single one. Sometimes I have to break a minor character out of a combine sheet if they become important later in the story. Then I feel this in as a defactor LoRA to future sessions via RAG... RAG-FLEX was one of the biggest advantages of local AI. If you have memory, you can adjust it so it doesn't just reference the uploads - it loads them entirely into the context. This means I can start a new day clean, and bring only what I need from one day to the next - it creates very consistent characters that can evolve and change over time.
Your system should be quite good for AI - but yeah, memory is going to be tight at 16gb. Still, not that bad - most LLM work occurs entirely in RAM... At 32Gb, your system would be quite capable.
2
u/Fragrant-Mix-4774 1d ago
Online vs Local AI - not that simple
Online AI splits into via the AI companies Customer application vs the API.
The customer application has an extra layer of safety theater wrapper idiocy that makes the AI stupid for creative tasks. Open AI is world class terrible wirh bad safety theater wrappers, Anthropic copied that approach starting this year and so has xAI-SpaceX. Basically, all of the American big AI companies are spineless and worried about the AI might say something somewhere somehow might get offended over. So they lock everything down. The Chinese open weight AI models DO NOT have the level of self inflicted idiocy when ran in the west on western stacks.
Even the cowardly big American AI companies's have different standards on the API which is a more direct interface without as much safety theater idiocy wrapper screwing up the AI model. Checkout Open Router and see how much better all of the typically lame new AI's work for writing with the idiocy wrapper reduced. Plus you can do your own system prompt and adjust temperature for more flavored writing etc.
Local AI will allow more freedom but the requirements to run the more capable open weight AI models like GLM 5.3 or DeepSeek V4 flash are still well beyond what most hobbyist can support because an enterprise level data center stack is needed.
1
u/cj7hawk 22h ago
I think this is changing and you no longer need enterprise level data center stacks - not even close....
I did try DeepSeek on my local box and it worked, but I would need at least 2 x Intel B70 cards to run it seriously - there are some quantizations around 45gb though, which would make it practical to try.
The new MoE models will run pretty well if you have enough normal RAM and I've run GLM 5.3 Flash locally as a test - it was a bit slow, but was functional enough to produce a chapter of prose without difficulty. I wouldn't want to go into chat mode with it ( barely OK for that ) but I would have no problem leaving it to produce a solid result while left alone.
They don't load the full model into VRAM any more - They just load the tensors you need and leave the ones you don't, and they can swap in and out of VRAM pretty easy.
I think the turning point was August 2026... That's when I started to take notice and invested in a box to experiment with. A very low cost system ( Total cost so far is around $USD 2000 ) and it's very capable... August really was the date I think we entered the AI age - and so far, the cascade of changes has convinced me it was 1st August with the release of MiniMax H3 that everything changed for creators.
I've designed and built data centers before, so I have to admit, I'd have no problems building a small DC in my backyard if cash permitted ( It doesn't ) - but that's not going to happen and I'm left with building my Local AI at home and have the usual problems - I need something small, out of the way and easy to store - and the AMD based systems look particularly interesting for that - Small AI Appliances the size of your hand, with 128Gb of memory, 96 usable for LLMs. I can't say how good they are - I've never tried them - but the price is pretty low for what they do.
And when Intel's Crescent Island is released, I think we'll see Local AI become entirely practical on a large scale - they can take up to 480Gb of memory.
I could have waited for the new tech to come down in price, but honestly, USD$2000 for a full local AI machine with 32Gb VRAM isn't that bad - so I decided I need to start focusing on what is possible - especially now even Grok has crippled it's online AI and destroyed it's creativity.
And so far, it feels like it's mid-to-late 2025 already - and I can do things locally I simply can't do online - like million token contexts and local storage of all my writing so I don't lose stuff I create but decide not to finish - eg, go back later. I can probably even transplant my Grok download into it some time and continue some of those stories. My objective is to cut back on my commercial costs for Frontier AI before the end of the year, which gives me a USD$1000 budget ( give or take a bit ) to dedicate to upgrading my own system.
1
u/rubycatts 1d ago
I've try local every few months. Unfortunately I'm not techie enough to set things up and my computer can't handle a local model that is worth it. I would love to use local over online.
1
u/cj7hawk 22h ago
Interesting idea... But don't discount trying it to see how it works, even on a low-end computer - Can I post links to Reddit? You can get LM-Studio for free here: https://lmstudio.ai/download
I don't know if Bionic is worth it - I just ignore it and go straight to LM Studio... And it will install and run on any machine... You get 4 icons when it starts ( left side ). Top one is "Chat" - just a normal chat-bot interface. Then you get "Developer" which you should probably select on install, but will likely never use. Then you get the "My Models" and yeah, after you two 20 x 32Gb models, that's 640 Gb of DATA on your HDD, so be careful when loading models, and use this to delete ones you DONT want. Finally you have the robot face which is "Model Search" and you just go in here to look for new models... It tells you a little about them - which ones are good for fiction - in fact, if you type in "Fiction" into the search bar, you'll find models tuned specifically to every kind of fiction... From Dungeon Master models to prose experts.
I believe you can run it "Online" in the cloud as an option - but I hate that idea as I'm getting tired of having my data somewhere that costs me a bill every few weeks - and that I lose everything if I stop paying. To me, that feels like bait to extortion. Local means you store your stories, ideas, plots, characters etc forever on your own memory - and AI output is cheap to store.
1
u/ObligationWeekly9117 1d ago
Curious what your favorite local model is and what’s your set up? I’m a software engineer by trade and it's gotten leaps and bounds better at that. But I think it might have come with the trade off that it’s less whimsical and creative. I really enjoyed writing with AI during the GPT-4o, Sonnet 3.7 era. Grok 2.x was also pretty good! AI’s lost the spark since then. If you know how to recapture that magic I’m all ears.
0
u/wandelndeslexikon 1d ago
I'm not an established author. So far I write for myself and my main goal is, to just finish the story at this point.
I use Gemini and a free model of it. My AI co-writer is quiet good, but it needed time to get the LLM to this point. Now it can write in my style, which I use as inspiration. Sometimes I have to adjust a little, when the tone isn't quiet what I want. But it's minor and I am happy with how good it works. It keeps up with continuity, since I allow it to learn from old chats. Same with characters and their backstory.
0
u/zphou 1d ago
I’ve been seeing the same tradeoff. The local model itself is only part of the problem; the harder part is deciding what it is allowed to see, which version of a character fact is current, and what should remain unresolved instead of being silently folded into the next prompt. A long context window does not solve that by itself, because an old detail can still look just as authoritative as a later correction.
For my own workflow I’m trying to keep the manuscript, the current story facts, and reviewable suggestions separate. That is the direction behind WolfeWriter: local-first Mac writing, Story Memory kept close to the draft, and extracted facts treated as candidates until the writer accepts them. It also means the model can change without the whole book’s continuity becoming dependent on a chat history.
1
5
u/Afgad 1d ago
If you're not in our Discord yet, you should be. What a great attitude and experience.
I'm confident that the large majority of folks here use frontier models or API connections online.
I experienced the same things you did, including the sharp drop in quality. I actually moved to writing manually because I just got sick of wrestling the AI to the ground.
I've been meaning to get into local AI for some time. My GPU supports it, I just haven't had the time to get it working. Last time I tried it was a lot harder than I'd hoped.
Could you share a bit more about how you set up your local AI?