Build on hardware you control.
Deployment, optimisation and apps for Apple Silicon, DGX / GB10 and RTX. Browse fixed-price services below, or discuss a larger private AI build.
Services and pricing
Set up your hardware
| Service | Turnaround | Price |
|---|---|---|
| Health check90-minute session on any box: thermals, network, memory headroom, drivers and containers. Written findings and a fix list. | 1 day | $250 |
| Two-box DGX Spark / GB10 clusterBoth ConnectX rails on real RDMA, one large model served across both boxes, benchmarked on your real prompts. Scripts and a runbook. | 2–4 days | $1,200 |
| Windows + RTX workstationDocker in WSL2 that stays up, a model server that survives updates and reboots, private access from your other machines. | 1–2 days | $500 |
| Apple Silicon Mac (MLX)A local model server your apps call like the OpenAI API, with image understanding. Runs offline. | 1 day | $400 |
Build private AI
| Service | Turnaround | Price |
|---|---|---|
| Coding agent on your own modelClaude Code, opencode or Hermes connected to your local model, tuned for speed and checked against the requests the agent really sends. | 1–2 days | $600 |
| Chat with your documentsA private assistant that answers from your files, wiki or tickets, with sources. Nothing is sent to a cloud API. | 3–5 days | $1,500 |
| Image and video studioImage generation and editing, short video with sound, on an RTX card or a Mac with MLX. Usable from a web form or as tools for your AI agent. | 3–5 days | $1,000 |
| Voice and audioPrivate transcription, text-to-speech and voice assistants with MLX Audio or on a GPU. Custom voices only from recordings you have the rights to. | 2–3 days | $800 |
Fine-tune models
| Service | Turnaround | Price |
|---|---|---|
| Fine-tune on your dataLoRA or QLoRA fine-tuning of a language model on your examples, on a DGX Spark, an RTX GPU or a Mac with MLX. Includes a before-and-after evaluation on held-out data. | 1–2 weeks | from $2,000 |
| Image style or product modelTrain an image model on your products, characters or brand style so it produces them consistently. | 1 week | from $1,200 |
| Convert and quantizeGet a model running on your hardware: MLX, GGUF, NVFP4 or int4, with the speed and accuracy measured against the original. | 2–3 days | $700 |
Advise and secure
| Service | Turnaround | Price |
|---|---|---|
| Which hardware should I buy?A written recommendation for your models and budget, from someone who runs DGX Spark, RTX and Mac setups every day. | 1 day | $300 |
| Security review of a self-hosted AIWho can reach your model server, what it exposes, how keys and data are handled. A prioritised fix list. | 2 days | $800 |
| Support retainerAsync help, model swaps and upgrades, with a backup and rollback plan before every change. | monthly | $300/mo |
Fixed quote before any work starts. Setup work includes a 14-day fix window. 50% up front, 50% on hand-off, or platform escrow. Larger projects quoted separately.
Private AI builds, scoped to your hardware.
For work beyond a fixed-price package, I validate the proposed stack on my own hardware before accepting the project. We agree on deliverables and acceptance criteria before implementation.
Apple Silicon & native AI apps
MLX on Mac, on-device text/vision/speech and SwiftUI apps for iPhone, iPad and MacBook.
Typical deliverable: A native app or model service checked on the intended device and workload.
DGX Spark / GB10 clusters
Two-box deployment, RDMA and NCCL diagnosis, distributed inference and startup optimisation.
Typical deliverable: A configured cluster, representative benchmarks and a runbook.
RTX workstations & creative studios
Language and vision serving, image/video workflows, memory budgeting and Docker / WSL2 reliability.
Typical deliverable: Working AI services fitted to your GPU and workflow.
Applications connected to private models
Coding agents, document assistants with citations, native interfaces and private API integration.
Typical deliverable: An integrated workflow tested from the application through the model server.
Model optimisation & ongoing operation
Model selection, quantization, targeted fine-tuning, endpoint security, monitoring and upgrades.
Typical deliverable: Before-and-after measurements, documented tradeoffs and a handover plan.
Security work covers technical assessment and mitigation; it does not include legal certification.
From first call to hand-off
Free 15-minute call
Your hardware, the model you want, and what you want it to do.
Fixed quote
You pick a service. No hourly surprises.
Access on your terms
A temporary SSH key or a shared screen. I never ask for passwords, and the key is removed at the end.
Snapshot first
Your current configuration is backed up before anything changes, so every step can be undone.
Hand-off
A runbook written for your setup, plus benchmark numbers on your own prompts.
FAQ
Why work with eAccelerate?
I operate Apple Silicon, two GB10 systems and an RTX workstation myself, and build the apps that use them. See my measured lab results and Apple app work.
NVIDIA's Cluster Assistant already connects two Sparks. Why pay?
It connects them. Whether NCCL is really on RDMA, both rails carry traffic, and the model reaches the recipe's speed on your prompts is a separate job.
Which models?
Open-weight language, vision, image, video and audio models: the Qwen, GLM, DeepSeek and Llama families, Qwen-Image, Whisper and others, at whatever precision fits your hardware.
Can fine-tuning make a small model as good as a big one?
For a narrow task, often close. For general knowledge, no. The evaluation before any fine-tune tells you which case you are in, before you pay for training.
Mac, Spark or RTX: which should I buy?
It depends on the model size and how fast you need answers. I run all three daily, so I can tell you what each one handles before you spend money. Software does not always carry over: engines built for Apple's MLX do not run on a Spark, and the reverse.
Do you host models for me?
No. Your hardware, your data. I don't resell inference.
Is this affiliated with NVIDIA or ASUS?
No. eAccelerate is independent consulting, based on running these boxes daily: two GB10s, an RTX PRO 5000 workstation and an Apple Silicon Mac running MLX.
Book a free 15-minute call
Email or message me with your hardware and what you want to run. I reply with a time for the call.