[K8s]Vertical Pod Autoscaler Isn't A Cloud Feature — I Put It On A Bare RKE2 Cluster And Its First Answer Was 10x Reality
Situation
Had Already Set Up VPA On A GKE Workload (A ClickHouse Backend, After It OOM’d Twice) And Assumed — Wrongly — That Vertical Pod Autoscaler Was Something GKE Provided As A Managed Feature. Turns Out The “Enable” Button On GKE Is Just Google’s Wrapper Around The Same Open-Source kubernetes/autoscaler Project Anyone Can Install Anywhere. Wanted To Confirm That For Real, So I Put It On A 6-Node Homelab RKE2 Cluster With No Cloud Provider In Sight.
It Went In Clean. What Wasn’t Clean Was The First Recommendation It Handed Back For A Live Workload — Ten Times The Container’s Actual Memory Usage, Within Minutes Of Being Installed.
Result First:
| Goal | Confirm VPA Works On Vanilla Self-Managed Kubernetes, Not Just Cloud-Managed Clusters |
| Prerequisite Check | Metrics Server — Already There. RKE2 Ships rke2-metrics-server As A Default Addon, Unlike Vanilla kubeadm Clusters |
| Install Method | Official In-Repo Helm Chart, Not The vpa-up.sh Shell Script — Skips Hand-Rolling Webhook Certs With openssl |
| Real Footprint | 6 Pods (2x Recommender / Updater / Admission-Controller), ~8m CPU + ~59Mi Memory Combined — A Rounding Error On 2-vCPU Nodes |
| The Surprise | Pointed It At A Real, Low-Stakes Workload (Headlamp). Actual Usage: 24Mi. VPA’s First Target: 250Mi |
Turns Out The Gap Isn’t A Bug — It’s What VPA Looks Like When It’s Being Deliberately, Structurally Cautious With A Number It Doesn’t Have Enough Data To Trust Yet. Which Is Its Own Kind Of Interesting, Once You Go Looking For Why.
Confirming VPA Is Cloud-Agnostic
The Claim To Verify: GKE’s --enable-vertical-pod-autoscaling Flag Isn’t A GKE-Native Feature — It’s A Managed Install Of The Same kubernetes/autoscaler Sub-Project (vertical-pod-autoscaler) That Ships On GitHub For Anyone To Run.
Two Prerequisites, Checked On The RKE2 Side Before Assuming Anything:
kubectl get pods -n kube-system | grep metric
# rke2-metrics-server-56b57b8dfb-ggnb9 1/1 Running
kubectl get apiservices | grep metrics
# v1beta1.metrics.k8s.io kube-system/rke2-metrics-server True
Metrics Server Was Already Running — RKE2 Bundles It By Default (Vanilla kubeadm Doesn’t, And That’s Usually The First Blocker People Hit Installing VPA On A DIY Cluster). One Fewer Thing To Set Up.
Two Install Paths, Picked The Less Manual One
The Official Repo Offers Both:
hack/vpa-up.sh— Clones The Whole Monorepo, Shells Out ToopensslTo Generate The Admission Webhook’s TLS Cert By Hand, Applies Raw Manifestscharts/vertical-pod-autoscaler— An Actual Helm Chart Living In The Same Repo, Where Cert Generation Is A Helm Hook Job Instead Of A Shell Script Step You Run Yourself
git clone --depth 1 https://github.com/kubernetes/autoscaler.git
helm install vpa autoscaler/vertical-pod-autoscaler/charts/vertical-pod-autoscaler \
--namespace vpa-system --create-namespace
Six Pods Come Up: Recommender x2, Updater x2, Admission-Controller x2 (Chart Default Replica Count, For HA). Combined Real-World Footprint After A Few Minutes Settled:
NAME CPU(cores) MEMORY(bytes)
vpa-...-admission-controller-cx99z 1m 10Mi
vpa-...-admission-controller-pjjlm 1m 10Mi
vpa-...-recommender-qsng2 2m 6Mi
vpa-...-recommender-zzgdv 1m 14Mi
vpa-...-updater-hn74p 1m 13Mi
vpa-...-updater-kww9b 2m 6Mi
~8m CPU, ~59Mi Memory, Total, Across All Six. On Nodes With 2 vCPUs Each, That’s Nothing. Whatever Hesitation I Had About Homelab Node Headroom Before Installing Turned Out To Be Unfounded.
Notes
- Installing VPA Does Nothing By Itself. The Components Sit Idle Until You Create An Actual
VerticalPodAutoscalerObject Pointing At A Workload. No Object, No Recommendation, No Resizing — Just Six Quiet Pods.
The Part I Didn’t Expect: The First Real Recommendation
Picked Headlamp (A Kubernetes Dashboard Deployment Already Running In The Cluster) As A Deliberately Low-Stakes Target — It Collects A Lot Of Cluster State Continuously, So It Seemed Like A Reasonable Thing To Watch. Attached A VPA Object In updateMode: "Off" — Recommendation-Only, Never Touches The Running Pod:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: headlamp-vpa
namespace: headlamp
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: headlamp
updatePolicy:
updateMode: "Off"
Checked Actual Usage Right Before Applying It:
NAME CPU(cores) MEMORY(bytes)
headlamp-fb548559f-zs8h2 2m 24Mi
Configured Resources On The Deployment At The Time: requests: 50m/64Mi, limits: 250m/256Mi. Already Comfortably Over-Provisioned Relative To Real Usage, By Eye.
Under A Minute After The VPA Object Landed:
status:
recommendation:
containerRecommendations:
- containerName: headlamp
lowerBound: { cpu: 25m, memory: 250Mi }
target: { cpu: 25m, memory: 250Mi }
upperBound: { cpu: 640m, memory: "2013Mi" }
Both Targets Land At Roughly 10x Actual Usage: 25m Against 2m Of CPU, 250Mi Against 24Mi Of Memory. Only One Of Them Bothered Me. 250Mi, Sitting Six Mebibytes Under The Container’s Existing 256Mi Limit, Looked Like Something That Needed Explaining, And An Upper Bound North Of 2Gi For A Dashboard App Idling At 24Mi Made It Look Worse. 25m Looked Harmless, And Was Even Lower Than The Container’s Current 50m Request, So I Read It As A Sane Safety Margin And Moved On.
Noting That Deliberately, Because It Matters Later: I Went Digging Into The Memory Number And Never Checked The CPU One, For No Better Reason Than That It Looked Reasonable.
Why Memory Recommendations Are Built To Run High
Went Digging Into The Recommender’s Own Source Rather Than Guess:
- Memory Targets Are Computed From The 90th Percentile Of Peak Usage Per Aggregation Window (Default 8 Days), Not Average Usage. CPU Is Compressible — Throttling A Container That Briefly Needs More CPU Just Slows It Down. Memory Isn’t — Exceed The Limit And The Kernel OOM-Kills The Container. The Recommender’s Math Is Deliberately Asymmetric For This Reason: It Would Rather Overestimate Memory Than Cause An OOM.
- With Very Little Observation History, The Confidence-Interval Math Pulls The Bounds Apart On Purpose, Never Toward A Smaller Number. The Lower Bound’s Confidence Multiplier Collapses Toward The Target Itself When History Is Thin (Effectively: “Don’t Suggest Shrinking Something You Haven’t Watched Long Enough To Trust”). The Upper Bound’s Multiplier Does The Opposite — Balloons Toward An Enormous Number When History Is Thin, Which Is Exactly The ~2Gi Ceiling Showing Up Here. Both Directions Of “I’m Not Sure Yet” Get Resolved The Same Way: Toward Not Under-Provisioning.
What That Doesn’t Fully Explain, And What I’m Deliberately Not Papering Over: Why The Target Itself Landed At Almost Exactly 250Mi — Suspiciously Close To The Container’s Existing 256Mi Limit — Rather Than Somewhere Nearer The 24Mi It’s Actually Using Right Now. Didn’t Find A Clean, Citable Answer To That One In The Time I Spent On It. Possible It’s An Artifact Of How The Very First Recommendation Gets Computed Before A Full Aggregation Window Has Elapsed; Possible It’s Something Else Entirely. Leaving It Genuinely Open Rather Than Inventing A Tidy Explanation That Sounds Right But Isn’t Verified.
Notes
- A Confident-Looking Number Isn’t The Same As A Confident Number. VPA Handed Back A Specific-Looking
250Mi, Not A Vague Range — But It Had Observed The Workload For Under A Minute. The Precision Of The Output Format Doesn’t Tell You Anything About How Much To Trust It Yet. - Memory And CPU Are Computed By Deliberately Different Maths. Memory Targets A High Percentile Of Peak Usage Specifically To Avoid OOMs; CPU Is Comparatively Relaxed Because Contention Only Throttles. Worth Knowing As Background — Though It Does Not, On Its Own, Explain The Specific Numbers Above.
- The Number You Don’t Question Is The One To Watch. Both Targets Here Are About 10x Actual Usage, But Only The Memory One Looked Alarming Enough To Investigate. That Asymmetry Is In The Reader, Not In The Data.
What Happens Next
This Is Deliberately Posted Before The Story’s Finished. The Whole Point Of Confidence-Weighted Recommendations Is That They’re Supposed To Get Better With More History — VPA’s Own LowConfidence Condition Exists Specifically To Flag When A Recommendation Shouldn’t Be Trusted Yet. Letting This Run Against Real (Light) Traffic, Then Coming Back To See Whether The Memory Target Actually Converges Down Toward The 24Mi It’s Really Using, Or Whether It Stays Anchored Near The Configured Limit Regardless Of Real Usage. And While I’m There, Checking The CPU Number I Waved Through. Part 2 Covers Both.
Real-World Application
| Scenario | What To Do |
|---|---|
| Deciding Whether VPA Is Worth Installing Outside A Cloud-Managed Cluster | It’s Fully Open-Source And Cloud-Agnostic — The Only Real Prerequisite Is Metrics Server. Check For It First; Some Distros (RKE2) Bundle It, Some Don’t |
| Installing VPA On Any Self-Managed Cluster | Use The In-Repo Helm Chart Over vpa-up.sh Unless You Have A Specific Reason Not To — Same Components, Less Manual Cert Handling |
| Reading A Brand-New VPA Recommendation | Check The RecommendationProvided / LowConfidence Conditions Before Acting On Any Number — A Recommendation Minutes Old Carries Very Different Weight Than One With Days Of History Behind It |
| A Recommendation That Looks Wrong Next To One That Looks Fine | Check Both. Memory And CPU Are Computed By Different Maths (Peak Percentile Versus A More Relaxed CPU Path), But That Doesn’t Mean The Reasonable-Looking Number Has Been Verified — It Means It Hasn’t Been Questioned |
Reference:
- kubernetes/autoscaler — Vertical Pod Autoscaler
- VPA Recommender Logic (source)
- Real Setup: Home RKE2 Cluster (6 Nodes), 2026-08-06 — Part 2 Pending Longer Observation Window