Apple Silicon · DGX / GB10 · RTX

Private AI.
On your own
hardware.

I deploy and optimise private AI on Apple Silicon, DGX / GB10 and RTX hardware—and build the native apps and workflows that make it useful.

Hands-on infrastructureTwo GB10 systems, a 48 GB RTX workstation and a Mac
Measured engineeringPerformance and quality checked against defined workloads
Practical handoverConfiguration, benchmarks and a runbook for your setup
Watch / ModelDeck in action

See the app. See the workflow.

A three-minute ModelDeck walkthrough, from the native interface to AI workflows. The demo shows the product experience; local and connected capabilities are explained on the Apple apps page.

ModelDeck · Product demonstration · 3 minutes Open video ↗
Selected engineering work / Personal lab

Three problems. Concrete results.

These case studies come from my own lab.

01 / Distributed inference

When a healthy cluster is still slow

103–116%
Measured on my cluster against published prefill references

Investigating a two-box setup whose diagnostics looked healthy while its interconnect underperformed.

Read the cluster case study →
02 / Operational performance

Getting a large model ready sooner

~13 → ~3 min
Measured time until ready for a first request

Reducing startup delay for a model spanning two machines, with accuracy checked after the change.

Read the startup case study →
03 / Multimodal infrastructure

One workstation, several AI workflows

300+ tok/s
Peak generation speed on my RTX PRO 5000

Language and vision serving, image editing and video generation on the same hardware platform.

Read the workstation case study →
Featured product / Apple platforms

AI in your hand.
And on your desk.

ModelDeck & Apple AI applications

Native app development for iPhone, iPad and MacBook, with on-device MLX inference and connections to private model infrastructure. A product experience that brings the model to the user.

iPhone

On-device AI

MLX text and vision, local speech workflows and memory-aware model lifecycle handling.

iPad

Native app experience

An Apple app experience for a larger screen, alongside access to private AI services.

MacBook

Apple Silicon AI

Mac applications and MLX models running on the MacBook, extending private AI to the desktop.

How I can help

The hardware. The models. The app.

Explore all services and fixed prices →

Apple Silicon & on-device AI

MLX models on Mac, on-device text and vision, and native apps for iPhone, iPad and MacBook.

DGX / GB10 clusters

Two-box deployment, RDMA and NCCL diagnosis, distributed inference and workload benchmarks.

RTX workstations

Language and vision serving, image/video workflows and reliable Docker / WSL2 operation within your memory budget.

Private AI applications

Coding agents, document assistants and native interfaces connected to models on hardware you control.

Optimisation & ongoing support

Model selection, quantization, targeted fine-tuning, upgrades and runbooks, measured against your actual workload.

One specialty / Three hardware platforms

Make the hardware you own work for you.

I work through deployment, memory limits, performance, application integration and recovery, so your private AI is usable from the devices and tools you already rely on.

Engagement approach

Define success before building.

01 / Assess

Match the workload to the hardware

Agree on users, data boundaries, hardware, budget and what a useful result looks like.

02 / Pilot

Test on real examples

Compare a baseline with the proposed approach using representative prompts and acceptance criteria.

03 / Deploy

Connect the workflow

Integrate the model with the intended tools, document access and back up configuration before changes.

04 / Improve

Measure and hand over

Check performance and output quality, deliver a runbook, and agree on ongoing support.

What delivery should include

Evidence you can inspect.

Scope and acceptance criteria are agreed before implementation. Measurements should describe the workload they actually tested.

  • Quality: representative examples, baseline comparisons and documented failure cases.
  • Performance: hardware, model, precision, prompt size and cold or warm timing conditions.
  • Privacy: agreed data flows and external dependencies, including any downloads or third-party services.
  • Operations: configuration backups, rollback instructions and a runbook for your environment.
About eAccelerate

Built from operating the systems myself.

eAccelerate is my independent consulting practice, focused on private AI running on client-owned hardware. My lab spans two GB10 systems, an RTX PRO 5000 workstation and Apple Silicon with MLX.

I build the applications as well as the infrastructure: native Apple apps, on-device MLX and private model services. For your project, I validate the proposed approach on my hardware before accepting implementation work.

Independent of NVIDIA and ASUS.

What you can count on

A clear scope. A careful handover.

Know the price first

A fixed quote and agreed deliverables before work starts. Setup work includes a 14-day fix window.

Keep control of access

A temporary SSH key or shared screen. I don’t ask for your passwords, and project access is removed at handover.

Leave with a working setup

Configuration backups before changes, a rollback plan, benchmarks on your prompts and a runbook written for your system.

Free / Five checks before you deploy

Is your private AI setup ready?

  1. Check the workload: define representative prompts and expected outputs before choosing a model.
  2. Check memory and cooling: measure under load, with the actual model and context size.
  3. Check the network path: for a cluster, verify the intended RDMA rails carry real traffic.
  4. Check the full request: time an application request, including model loading and first-token latency.
  5. Check recovery and access: test restart and rollback, and review who can reach the endpoint.
Download the free five-check list →
Learn / Guides

Go deeper into the engineering.

Explore my existing resources on training an LLM, environment setup, fine-tuning and MLX projects.

Start with your problem

What do you want your AI system to do?

Email or message me with your hardware, intended workflow and the problem you want to solve. We’ll arrange a free 15-minute call and agree on scope before any work starts.