Multimodal infrastructure / Personal lab
One workstation, several AI workflows
EnvironmentRTX PRO 5000 / 48 GB
FocusLanguage, image and video
MeasurementMixed cold / warm timings
The problem
I wanted language, vision, image and video workflows on one RTX PRO 5000 workstation with 48 GB of GPU memory. The challenge was fitting the serving stacks and model-switching workflow to the available resources.
What I built
My language and vision endpoint uses Qwen3.8-27B on SGLang with NVFP4 weights and a DFlash2 speculative drafter. Qwen-Image handles image generation and editing, alongside a local video workflow with audio.
Measured language performance
I measured peak generation above 300 tokens per second, prompt processing at 6,073 tokens per second on a 22,000-token prompt, and 3.7 seconds to the first token on that prompt. The endpoint supports a 262,144-token context window; the timing here is for the specified 22,000-token workload.
Measured creative workflows
A 1536×864 image took 18 seconds, an image edit 8 seconds and a 1024×1024 transparent PNG 26 seconds. A five-second video with sound at 864×480 took 34 seconds. Image times include cold startup; video time is warm. These timings describe separate requests, rather than simultaneous generation.
What this means for your system
I evaluate the complete user workflow, including model loading, generation and output delivery. The right serving and switching plan depends on your memory budget and the workloads you need.
Case study from my own lab. Results depend on the hardware, model and workload.