AIAI Tech Engine
August 2026 · only machines on our desk

A desk-side local AI lab
16GB CUDA vs 64/96GB Mac vs 128GB Spark on site

We only write what we have checked on the four machines we hold. Unmeasured tok/s, other people's card benches, and nameplate scores for hardware that has not arrived are not our results. The Gigabyte DGX Spark 128GB arrived on 2026-08-31 and is on site. We still have no tok/s for it. Strix Halo is still planned.

Current fleet
4 machines
M1 Max 64GB · M2 Max 96GB · RTX 4080 16GB · Spark 128GB
Still planned
Strix Halo
AMD AI Max+ 395 128GB · not arrived
27B Q4 today
Mac first
4080 16GB is tight · Spark tok/s empty
Unmeasured
tok/s blank
Published only after we time the same prompt
SPONSORED / ADVERTISEMENT
Reserved ad slot — we only wire a unit ID issued by the AdSense console after approval.

Lab notes & analysis

Checked on hardware we hold. Speed tables only after a measurement.

9 reports published
Qwen3.8-Flash-Next2026-08-31AI Tech Engine Lab

Qwen3.8-Flash-Next memory tiers: what fits in 16GB / 64·96GB / planned 128GB

Reading the 2026-08-26 Qwen3.8-Flash-Next (125B-A6B + 51B n-gram) from Unsloth GGUF, community MLX file sizes, and the official card only. Mapped onto the AI Tech Engine lab’s RTX 4080 16GB / M1 Max 64GB / M2 Max 96GB / planned 128GB. No unmeasured tok/s.

#Qwen3.8-Flash-Next#local LLM#memory tiers#llama.cpp#MLX#testbed
Read the report →
Qwen3.82026-08-30AI Tech Engine Lab

M1 Max 64GB timings: Qwen3.8-27B (MTPLX · omlx) vs Ornith-1.5-35B

On the same MacBook Pro (M1 Max 64GB) we run Qwen3.8-27B through MTPLX and omlx DFlash, and Ornith-1.5-35B-A3B through omlx, on the same needle and the same tokens. All three paths hit the needle. The same Angry-Birds-style HTML is playable only from Ornith. We do not use other people’s card benches.

#Qwen3.8#Ornith#MTPLX#omlx#M1 Max#local LLM#long context#measured
Read the report →
local LLM2026-08-30AI Tech Engine Lab

The 2026 local AI lab: 16GB CUDA vs 64/96GB Mac vs planned 128GB boxes

How the AI Tech Engine lab is actually wired. What the RTX 4080 16GB, M1 Max 64GB, and M2 Max 96GB do today, and why the planned Strix Halo and DGX Spark 128GB boxes are for 70B-class work and local agents. No unmeasured TPS.

#local LLM#hardware#testbed#Apple Silicon#RTX4080#DGX Spark#Strix Halo
Read the report →
Qwen3.82026-08-30AI Tech Engine Lab

Qwen3.8-27B and August 2026 open weights: what this lab can load today

We check Qwen3.8-27B and Meta Muse Glimmer 30B against Hugging Face and official docs, then map them onto the AI Tech Engine lab’s M1 Max 64GB / M2 Max 96GB / RTX 4080 16GB and the planned 128GB boxes. No unmeasured tok/s.

#Qwen3.8#Muse Glimmer#open weights#local LLM#Apache-2.0#testbed
Read the report →
hardware2026-02-23AI Tech Engine Lab

2026 local-LLM PCs that do not waste the budget: how to pick a GPU

RTX 4090, 4080 Super, used dual 3090s, Mac Studio M3 Max. What actually fits at each VRAM size, and which builds make sense at each budget — a buying map, not a lab fleet list.

#hardware#GPU#local LLM#RTX4090#VRAM
Read the report →
Cloudflare2026-02-23AI Tech Engine Lab

A $0/month serverless AI media stack (Astro + Cloudflare Pages)

Skip WordPress and a fat Vercel bill. Pair a static site generator with Cloudflare Pages and serve a tech blog from 300+ edge cities in well under a second.

#Cloudflare#serverless#Astro#web infra#monetization
Read the report →