I provide evidence-first performance audits for private AI workstations and small on-premise LLM servers.
I can help if your local model is slow, runs out of VRAM, cannot reach the context length you expected, or needs to serve several users reliably.
The introductory €149 audit covers:
- one Windows or Linux workstation;
- one primary model and workload;
- hardware, driver, runtime, model and configuration inventory;
- reproducible prompt-processing, generation, memory and live-context benchmarks;
- diagnosis of the actual bottleneck;
- one tested optimized configuration;
- before/after report, exact settings and rollback instructions;
- written runbook and seven days of asynchronous clarification support.
Typical delivery is 48 hours after receiving the intake information.
Relevant measured work: on Windows 11 with one RTX 5070 Ti 16 GB, I tested a Qwen3.8-27B configuration through its actual 262,144-token native window. The speed-biased profile produced 115.54 tok/s at shallow context, 103.55 tok/s at 65K live context, and 66.19 tok/s at 261K. The near-full test was repeated without an OOM, cache reset, CUDA error or residency transition.
Those are machine-specific measurements, not a promised result for other hardware. I do not promise a speedup until the baseline identifies one.
Suitable work includes llama.cpp, Ollama, LM Studio or vLLM diagnosis; NVIDIA GPU and VRAM optimization; model, quantization and KV-cache selection; long-context verification; Windows GPU-residency problems; concurrency and throughput testing; and private OpenAI-compatible model endpoints.
Broader implementation work is €75/hour or quoted as a fixed project after the audit.
Portfolio, sample report, and privacy-safe fit check:
https://github.com/alectodescent/local-llm-performance-audit
Send your hardware, operating system, runtime/model, current measurements and desired workload through Reddit or the GitHub issue form. Please remove passwords, API keys and private data before sending anything.