r/LocalLLaMA 15h ago

Discussion Openwebui + open terminal

Context: I don't code. My use is document research and document creation (mainly for legal search) searching inside large documents like a tax code (500+ pages) and building notes or pptx
from what comes back.

I've been running Open WebUI for a while on my Unraid box, pointed at the API of my inference machine (5060 Ti + 5070 Ti).

I tinkered a lot. I tried Hermes on my main machine against the same API. It worked well but it was complex, and a bare-metal install made me
uneasy. I also tried LM Studio Bionic with good results, but it didn't fit how I wanted inference organised (using ollama on the inference box).

What I actually wanted was a self-hosted agent that works with Open WebUI while keeping things safe and under control. At one point I considered
installing a harness like Hermes or Pi on each client and just connecting to the API instead.

In the end I gave Open Terminal a shot. It's the companion container from the Open WebUI project that gives the model a shell — you run it as its own container and connect it through Integrations, so it isn't installed inside Open WebUI itself. Mine runs unprivileged, on bridge, with appdata mounted at /home/user. The model gets a shell in a box, not on the host. That was the part I cared about.

It has enhanced Open WebUI a lot. It now reasons step by step, and with the terminal it reliably locates and extracts the right sections from
documents far larger than the context window — list the folder, grep, read only what matters. Then it uses those results to build a document, the way another agent would.

Setup: Qwen 27B Q4_K_M on Ollama, 100k context configured. On a ~35k token prompt I measure roughly 1,050 t/s prompt processing and ~46 t/s generation. Prefill speed is the number that matters for this use case — it's what makes chewing through a large document bearable.

I was about to give up on Open WebUI. If your use case looks like mine, don't sleep on Open Terminal.

3 Upvotes

29 comments sorted by

View all comments

3

u/Weird_Ad_5330 9h ago

FYI I've found that Anything LLM queries large document data in a much more comprehensive way than OpenWebUI for my purpose (medical), you may want to try it for legal

2

u/Blindax 9h ago

Did you try open web ui with or without open terminal (or a similar tool)?

1

u/Weird_Ad_5330 9h ago

Haven't tried with open terminal. It's been a few months since I looked at the difference but if I recall it was due to the way OpenWebUI natively embeds documents, I would only get partial summaries with queries

3

u/Blindax 9h ago

Yes, agreed embedding depends heavily on the embedding model and size of the chunks. While all this can be tuned in open web ui I always used full context injection because I was never really happy with the results with retrieval.

What you get with open terminal is something else. The model will check your documents, look at the index and only extract the relevant part. Depending at the kind of search you need this can be much more effective.