[AI] LM Studio Couldn't Run This Model, So I Got It Running On My Own Mac
Situation
Been Looking At New Models Lately, And Came Across Muse Glimmer, A 30B Model That’s Supposed To Be Really Good For Personal Use. Wanted To See If It Could Replace Gemma-4 As My Log-Summarizer Backend, Plus A Similar Role On My Personal Assistant. LM Studio Was Already Running On This Machine, So I Just Tried Loading It There. It Failed Immediately:
No module named 'mlx_vlm.speculative.drafters.muse_glimmer'
After Actually Checking, LM Studio’s Bundled MLX Runtime Was Already The Latest On Both Stable And Beta Channels. There Was No Update Coming To Fix This. The Model Was Just Too New For What LM Studio Was Shipping.
Result First:
| Trigger | Wanted To Test Whether A New 30B MLX Model Could Replace Two Existing Local Models |
| Root Cause | LM Studio’s Bundled MLX Runtime Doesn’t Support The New Architecture Yet, No Update Available |
| Fix | Raw mlx_vlm Pip Package’s In-Process API, Wrapped In A Small OpenAI-Compatible Shim Server |
| Outcome | Model Now Serves Both Consumers Through The Same /v1/chat/completions Contract They Already Spoke To LM Studio With |
The Runtime Gap Turned Out To Be The Easy Part. The Real Surprise Was Finding Out MLX Itself Can’t Run Anywhere Except On The Machine In Front Of Me.
How It Broke (Root Cause)
The Runtime Is Frozen, The Model Isn’t
LM Studio Ships Its Own MLX Runtime As A Versioned Package (mlx-llm-mac-arm64-apple-metal-advsimd, Currently 1.11.0). New Model Architectures On Hugging Face Show Up Faster Than LM Studio Can Ship Runtime Updates For Them. Muse Glimmer Uses A Speculative-Decoding Module Layout The Bundled Runtime’s Loader Simply Doesn’t Know About. Not A Config Problem, Not A Quantization Mismatch, Just Code That Isn’t There Yet.
I Confirmed This Wasn’t Something I Could Fix By Updating LM Studio: Checked lms runtime ls Against Both The Stable And Beta Channels, Same Version On Both. Nothing To Update To.
Wanting To Centralize It Ran Into A Harder Wall
My First Instinct Was To Run This As A Service On My Home Kubernetes Cluster (rke2-home) So It’d Be Managed The Same Way As Everything Else. Checked The Node List:
rke-m01 / rke-n01–n05 all Ubuntu 20.04 x86 Linux
No Apple Silicon Anywhere. MLX Is An Apple Silicon + Metal Framework. It Doesn’t Run On Linux At All, And Critically, It Doesn’t Run Inside A Docker Container Even On A Mac, Because Docker Desktop For Mac Runs A Linux VM Underneath With No Metal GPU Passthrough. This Isn’t A Performance Tradeoff You Can Configure Around. If You Want MLX, It Runs As A Native macOS Process On Apple Silicon, Full Stop.
That Ruled Out The Tidy Kubernetes Answer. The Model Had To Live On The Mac, Which Meant The “Service” Part Had To Come From macOS’s Own Tools.
The Fix
Skip The Runtime, Use The Library Directly
The Raw mlx_vlm Pip Package (0.7.1) Supports Architectures The Bundled LM Studio Runtime Doesn’t. It’s The Same Underlying Library LM Studio Uses, Just Not Frozen At An Old Version. Its In-Process API Is Exactly What I Needed:
from mlx_vlm import load, generate
model, processor = load("mlx-community/Muse-Glimmer-30B-4bit") # load once
result = generate(model, processor, prompt, max_tokens=4096, temperature=0.2)
load() Only Needs To Run Once. Everything After That Is A generate() Call Against An Already-Resident Model, Which Is The Whole Trick For Turning This Into A Server Instead Of A CLI That Reloads 19GB Of Weights Every Invocation.
Wrap It So Nothing Downstream Has To Change
Both Consumers (A Go Service And A Node-Based Personal Assistant) Already Spoke To LM Studio Through The Same Contract: POST /v1/chat/completions, OpenAI Request/Response Shape. So The Shim Just Had To Speak That Same Contract:
@app.post("/v1/chat/completions")
async def chat_completions(request: Request):
body = await request.json()
history = to_mlx_history(body["messages"]) # OpenAI shape -> mlx_vlm's chat history shape
prompt = apply_chat_template(processor, model.config, history, num_images=0)
result = await anyio.to_thread.run_sync(lambda: generate(
model, processor, prompt,
max_tokens=body.get("max_tokens", 4096),
temperature=body.get("temperature", 0.2),
))
return {"choices": [{"message": {"role": "assistant", "content": result.text}}]}
Neither Client Needed A Single Line Changed. They Just Point At A Different Port.
Make It Actually Stay Up
launchd Handles The “This Needs To Survive A Crash Or A Reboot” Part macOS-Natively. No Docker, No Extra Layer:
<key>KeepAlive</key><true/>
<key>RunAtLoad</key><true/>
Killed The Process With kill -9 To Check This Actually Works Instead Of Trusting The Plist. New PID Came Up Within Seconds, Model Reloaded, /health Went Green Again.
One Serialization Detail That Matters
MLX Generation Isn’t Safe To Call Concurrently Against The Same Model State. A Single threading.Lock Around The generate() Call Keeps Requests Queued Instead Of Corrupting Each Other. That’s Fine For A Single-User Local Shim, But It Would Need Real Batching For Anything With More Than One Caller At A Time.
A Week Later: Staying Up Isn’t The Same As Staying Fast
KeepAlive Solved The Crash Case. It Does Nothing For The Case Where The Process Is Alive But Barely Moving.
A Week After Setup, I Noticed Every Alert Summary Had Quietly Moved To The Cloud Fallback. From The Outside The Shim Looked Fine: /v1/models Answered In Under 0.1 Seconds, So The Health Probe Passed Every Time. The Log Told A Different Story:
2026-09-22 16:13:40 INFO generated 1053 tokens in 162.3s (14.6 tok/s), prompt_tokens=573
2026-09-25 17:25:35 INFO generated 145 tokens in 10867.7s (6.3 tok/s), prompt_tokens=23484
2026-09-27 05:19:04 INFO generated 966 tokens in 6000.9s (0.8 tok/s), prompt_tokens=387
Same Model, Same Machine, 10 To 20 Times Slower. One Request Took 100 Minutes, And Earlier A Single 23K-Token Prompt Held The Generation Lock For Three Hours, So Anything That Arrived Behind It Had To Wait Too.
The Most Likely Cause Is Memory Pressure. This Is A 32GB Mac Holding A Roughly 19GB Model Next To Everything Else I Run, And Swap Was Sitting At 6.7GB Of 8GB. That’s As Far As I Can Prove It. What I Did Confirm Is That A Restart Brought It Straight Back To About 15 Tokens/Sec.
And While Retesting, I Tripped Over My Own Note Below: Sent max_tokens=80, Got An Empty Reply, And For A Minute Thought The Shim Was Broken. The Reasoning Had Simply Eaten The Whole Budget.
Notes
- A Runtime Version Being “Latest” Doesn’t Mean It Supports The Latest Models. LM Studio’s Runtime And The Models You Can Download Through Hugging Face Move On Independent Schedules. When Loading A Brand-New Architecture Fails With A Python Import Error Rather Than A Normal LM Studio Error Dialog, That’s The Tell: It’s Not Rejecting The Model, It Genuinely Doesn’t Have The Code Path.
- MLX Only Exists On The Metal Side Of The Fence. Not Slower On Linux, Not Emulatable, Not Containerizable On The Same Mac It’s Running On. Anyone Planning To Put An MLX Workload Behind Kubernetes Needs To Know This Before Drawing The Architecture Diagram, Not After.
- The In-Process API Was Already There, Just Not Documented As “Build A Server With This.” Every Example I Found Was CLI-Shaped (
python -m mlx_vlm.generate), Which Reloads The Model Every Call. Theload()/generate()Split Underneath That CLI Is What Makes A Persistent Server Possible, And It’s One Import Away. - First Quality Test Looked Bad For The Wrong Reason. Ran It Once With
max_tokens=50And Greedy Decoding And Got A Leaked Internal Channel Token Plus A Truncated Reasoning Trace. That Wasn’t The Model. It’s A Reasoning Model, And 50 Tokens Isn’t Enough Room For It To Finish Thinking Before The Budget Runs Out. Bumping Tomax_tokens=4096Andtemperature=0.2With A Realistic Prompt Produced Clean, Correctly-Structured Output Every Time After. - The Real End-To-End Test Is The One That Matters, Not The Curl Test. Manually Curling The Shim Looked Perfect. Only Firing An Actual Alert Through The Full Production Pipeline (Log Query → Retrieval → This Shim → Escalation Logic) Confirmed It Behaved Correctly Under Real Conditions, Including Correctly Recognizing When It Had No Real Evidence To Work With And Asking To Escalate Instead Of Guessing. Generation Took 98 Seconds For 1,125 Tokens On This Machine (~14.5 Tokens/Sec For A 30B Model), Which Is The Kind Of Number You Only Get By Actually Running It Against Real Load, Not A Short Test Prompt.
- A Health Check That Answers Isn’t A Model That Works.
/v1/modelsProves The Process Is Alive, Not That It Can Generate At A Usable Speed. Tokens/Sec Per Request Is Already In The Log, And That’s The Number Worth Watching.
How To Prevent It
| Scenario | What To Do |
|---|---|
| A New Model Fails To Load In LM Studio With A Python Import Error | Check lms runtime ls Against Both Channels First. If It’s Already Latest, The Runtime Genuinely Doesn’t Support The Architecture Yet, Waiting Won’t Fix It |
| You Need An MLX Model Reachable By Multiple Services | Wrap It In A Minimal OpenAI-Compatible Shim With load() Once + generate() Per Request, Not A CLI Invocation Per Call |
| You’re Tempted To Put MLX Behind Kubernetes Or Docker | Don’t. It’s Apple Silicon + Metal Native Only, No Container Path Exists, Plan For A launchd-Managed Native Process Instead |
| You’re Judging Model Quality From A Short Test | Use A Realistic Prompt And Enough max_tokens For The Model To Actually Finish. A Truncated Reasoning Model’s Output Tells You Nothing About Its Real Quality |
| A Local Model Endpoint Passes Health Checks But Downstream Keeps Falling Back | Look At Per-Request Tokens/Sec In The Shim Log, Not The Health Endpoint. Cap Prompt Size So One Huge Request Can’t Hold The Lock For Hours, And Watch Swap On The Host |
Reference: