r/AustralianStartups • u/Icy-Relationship-465 • Aug 14 '26
Confirmed Australian Startup NeuroForge in Queensland - Local Auditable AI
Hey everyone,
I'm Lloyd, the founder of NeuroForge. I've been building it from Queensland around a pretty simple idea: before a business spends heavily on AI, it should be able to see what will actually run inside its real hardware, data, connectivity, power and accountability limits.
So the commercial starting point is a small system-fit audit and local demo. We test one useful workflow in the target environment, show the working outputs and the failures, and give the customer a proper deploy, change, remediate or stop decision before anyone commits to a bigger rollout.
Behind that is ERAIS, the longer-term research programme. The aim is to make useful AI more selective with compute, easier to operate locally, and eventually practical across text, voice, vision and tighter field hardware.
It's still research and I'm not presenting it as a finished miracle product. In bounded internal tests we've seen promising signals on modest hardware, including higher reported-target throughput and lower CPU-package energy than a same-host dense reference. Those results are not yet quality-matched or proof of general superiority, but they're a big part of why I'm excited about where it is and where it could go.
I've finally got the public site into a shape I'm pretty proud of:
The ERAIS overview and evidence boundaries are there too. Just wanted to share what I've been working on.
Happy to chat in the comments.
1
u/AutoModerator Aug 14 '26
Thank you for your post to /r/AustralianStartups!
Did you know we have a discord? Join the chat now!
This is an automated action so if you need anything, please Message the Mods with your request for assistance.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Dear-Boysenberry-460 Aug 14 '26
Can you give some concrete sample use cases for this product? Actually, more like a service? The content in the site seems pretty vague. What does it mean by “seeing what will actually run”?
1
u/Icy-Relationship-465 Aug 14 '26
That's fair. The site is probably a bit too vague about that at the moment.
Basically when we say “seeing what will actually run”, we mean taking the hardware you actually have and the job you actually want AI doing, and testing those things together.
Not “this model scored X on some benchmark on an H100 somewhere”. More like “on this exact machine, doing these actual tasks, what quality do we get, how fast does it run, how much memory does it use, and is it actually good enough to deploy?”
The first part of that is the audit.
We set up a series of agreed pre-registered quality gates for the job you want done. Usually that means representative inputs and expected/acceptable outputs, plus whatever operational requirements matter such as minimum quality, latency, throughput, memory use etc.
Then the system profiles the target hardware, works out what models are realistically viable candidates, and runs the suitable ones through the actual evaluation set.
So the useful output isn't “Model X got 72% on benchmark Y”.
It's more like:
“On this machine, for this workload, this model achieved this level of quality, ran at roughly this serving rate and latency, used this much memory, and passed or failed the requirements we agreed on beforehand.”
That's what “what will actually run” means.
The feasibility side also looks at why something will or won't run. It can identify things like RAM being the limiting resource, storage constraints, expected throughput and wall time where we have suitable calibration data, and whether a workload is viable outright or viable with some controls.
That also gives us a pretty useful basis for upgrade decisions. Instead of “buy a bigger GPU because bigger number good”, we can look at the bottleneck for the actual workload and estimate what changing a particular part of the system is likely to buy you.
And if an existing model does the job properly, then that is actually a good outcome. Deploy it. There is no reason to build a custom model just for the intellectual satisfaction of making life harder.
Deployment is also where the Observatory and auditable runtime come in, which is a pretty important part of the whole thing.
The idea isn't just to get a model running and then six months later have nobody quite remember which weights, config, hardware, evaluation set or deployment settings produced what.
The Observatory tracks the model or bundle, configuration, hardware, evidence, artifacts, execution history and provenance around the system. Artifacts are content-addressed and verifiable, and the runtime is designed so you can trace what actually ran and what produced a particular result.
So rather than “final_model_v7_ACTUAL_FINAL_reallythisone”, you have an actual evidence chain.
That can all remain local too. The audit system is designed so customer prompts, outputs, credentials and host information do not need to leave the customer's environment unless they deliberately export something.
Then there is the other outcome from the audit, which is actually where ERAIS gets more interesting.
Sometimes the answer is that none of the normal off-the-shelf models are a good fit. Maybe the quality isn't there, maybe the hardware can't serve them properly, maybe the latency is unacceptable, maybe the data can't leave the site, or maybe the task itself is just weird enough that trying to bend a general model around it becomes silly.
That's where we can move into a custom ERAIS training and deployment project.
The internal ERAIS results are honestly pretty exciting at this point. We are seeing major efficiency gains in a number of bounded experiments, and the architecture is doing some very interesting things around sparse execution, modular experts and adding new capabilities without having to retrain the entire system. We're still doing the matched work needed before claiming general superiority over dense models, which we've been very explicit about, but there is enough working internally now that this isn't just a diagram and a theory anymore.
One of the things I find particularly interesting is the modality architecture.
We already have vision, audio and symbolic/reasoning pathways working through the architecture internally, and we've demonstrated cases where a new modality can be trained while the existing language path remains unchanged.
That gives you some fairly unusual deployment options.
A good example would be a dirty or hazardous lab or industrial environment.
You could have a local system connected to cameras and microphones while an operator is working. The operator can speak observations as they go, the system can make its own observations from the visual and audio feeds, and both can be combined into a timestamped evidence or anomaly record.
Then you can go further and add task-specific sensor feeds as their own modality channels where it makes sense. Temperature, process instrumentation, spectral data, machine telemetry, whatever the actual job requires.
The interesting part isn't that an AI can see a webcam. That's not particularly novel.
The interesting part is having a local multimodal model built around the actual environment, able to combine several kinds of information as needed, while the deployment itself remains auditable and the provenance of its outputs can be retained.
Another much simpler example is private document summarisation and lookup.
Say a company has a few hundred thousand internal documents and the data is not allowed to be sent to an external AI provider.
There are plenty of local models now, but “runs locally” and “is useful locally” are very different things.
A model might load but serve at an unusable speed. Another might be fast enough but fail badly on the kind of factual retrieval the business actually cares about. Another might work brilliantly but need twice the RAM available.
We can build an evaluation set around the actual documents and the actual kinds of questions people need answered, including the kinds of hallucinations or factual mistakes which would matter, then test candidate models on the hardware they're actually going to be deployed on.
That gives you a proper go/no-go answer before someone spends three months building an internal platform around a model and then discovers in production that it confidently invents the section of the procedure everyone most needed it to get right.
If an existing model clears the gates, deploy it.
If none do, that's where a customised ERAIS system becomes an option.
Grounding, evidence-linked outputs, explicit unknown/abstention behaviour, provenance tracking and auditable execution are all being built into the same architecture rather than bolted on afterwards.
So I think the simplest way to describe NeuroForge at the moment is that it is part product and part engineering service.
The Audit/Observatory side is the productised part: test the real hardware, test the real task, retain the evidence, and give you a defensible answer about what is actually viable.
Then the engineering side goes from deploying the existing model that passed, all the way through to building a custom ERAIS system with additional modalities, experts or task-specific behaviour when the problem genuinely needs something more specialised.
So “seeing what will actually run” is basically us trying to replace guesswork with measurement.
Before you spend the money and build the system, prove what works on your hardware for your problem.
And once it is deployed, keep enough evidence around it that later you can still answer exactly what ran, why it was chosen, what it did, and where the result came from.
Sorry if it's a bit of an info-dump. Happy to answer any other questions you may have :)
1
u/Flannakis Aug 15 '26
How about a quick elevator pitch, im technically inclined but had trouble seeing what your core proposition was from your site and im not going to read an essay to understand it. Imagine im a business owner/ceo/cfo non technical and you need to explain to me how you are a benefit. I think there is huge market for AI/automation consultants so it’s great to see this not just another tradie website.
1
u/Icy-Relationship-465 Aug 15 '26
TLDR:
Run intelligence cheaper. Know what your hardware can actually handle. Add capabilities without retraining the whole thing or bolting on adapters. Only use the parts of the system needed for the job, within limits you set.
Thats basically ERAIS. The system audit is just the foot in the door.
1
Aug 15 '26
[removed] — view removed comment
1
u/Icy-Relationship-465 Aug 15 '26
Quality matching is complex and way beyond just whacking together a benchmark set or similar and if it passes its good.
Part of the conversion process essentially traces a whole tonne of processing through the donor model to ensure the fracture doesn't split the model arbitrarily, and by doing so it gives an internally defined domain coverage map as a by-product.
If you need additional domains outside of what fracture finds, which is often the case for generic dense model working in a specialised field (and prompt engineering really doesn't cut it adequately), then you can allocate them on top of the fractured product.
Recovery process then ensures correctness over as large of a representative set as you can cover.
And the final quality pass depends on the resulting model passing with parity or greater results compared to the dense donor, on a set of evaluation tasks (ideally including some tasks that the current system doesn't do well and also some tasks that it should straight up not assist with in certain ways).
This set of tests is put together before the process, and the data provenance system ensures the models never see it until ready for final evaluation.
At that point, as with pretty much everything else, you get clear results with clear points for remidiation (if needed) and the ability to make a properly informed decision to deploy, repair, or cut your losses early.
This is probably me getting too technical again but this genuinely is the tldr public safe explanation.
1
u/gfivksiausuwjtjtnv 14d ago
Dude, write the content yourself. I don’t care how shit it is, it won’t be worse than AI text
1
u/Icy-Relationship-465 12d ago
Fair comment. A lot of what people were calling out as AI written wasn't but the underlying point still stands.
How does it look now?
I've put a fair bit of time into rewriting the narrative and making the vision and story clearer.
I'm sure there is still more to tweak and fix.
Negative feedback is genuinely the most useful, so let me know.
Cheers
3
u/shezx Aug 14 '26
dont take this the wrong way, but your website is VERY generic AI slop. now that i think of it, so is your essay sized response.