Situation

Had Already Set Up VPA On A GKE Workload (A ClickHouse Backend, After It OOM’d Twice) And Assumed — Wrongly — That Vertical Pod Autoscaler Was Something GKE Provided As A Managed Feature. Turns Out The “Enable” Button On GKE Is Just Google’s Wrapper Around The Same Open-Source kubernetes/autoscaler Project Anyone Can Install Anywhere. Wanted To Confirm That For Real, So I Put It On A 6-Node Homelab RKE2 Cluster With No Cloud Provider In Sight.

It Went In Clean. What Wasn’t Clean Was The First Recommendation It Handed Back For A Live Workload — Ten Times The Container’s Actual Memory Usage, Within Minutes Of Being Installed.

Result First:

Goal Confirm VPA Works On Vanilla Self-Managed Kubernetes, Not Just Cloud-Managed Clusters
Prerequisite Check Metrics Server — Already There. RKE2 Ships rke2-metrics-server As A Default Addon, Unlike Vanilla kubeadm Clusters
Install Method Official In-Repo Helm Chart, Not The vpa-up.sh Shell Script — Skips Hand-Rolling Webhook Certs With openssl
Real Footprint 6 Pods (2x Recommender / Updater / Admission-Controller), ~8m CPU + ~59Mi Memory Combined — A Rounding Error On 2-vCPU Nodes
The Surprise Pointed It At A Real, Low-Stakes Workload (Headlamp). Actual Usage: 24Mi. VPA’s First Target: 250Mi

Turns Out The Gap Isn’t A Bug — It’s What VPA Looks Like When It’s Being Deliberately, Structurally Cautious With A Number It Doesn’t Have Enough Data To Trust Yet. Which Is Its Own Kind Of Interesting, Once You Go Looking For Why.

Confirming VPA Is Cloud-Agnostic

The Claim To Verify: GKE’s --enable-vertical-pod-autoscaling Flag Isn’t A GKE-Native Feature — It’s A Managed Install Of The Same kubernetes/autoscaler Sub-Project (vertical-pod-autoscaler) That Ships On GitHub For Anyone To Run.

Two Prerequisites, Checked On The RKE2 Side Before Assuming Anything:

kubectl get pods -n kube-system | grep metric
# rke2-metrics-server-56b57b8dfb-ggnb9   1/1   Running

kubectl get apiservices | grep metrics
# v1beta1.metrics.k8s.io   kube-system/rke2-metrics-server   True

Metrics Server Was Already Running — RKE2 Bundles It By Default (Vanilla kubeadm Doesn’t, And That’s Usually The First Blocker People Hit Installing VPA On A DIY Cluster). One Fewer Thing To Set Up.

Two Install Paths, Picked The Less Manual One

The Official Repo Offers Both:

  1. hack/vpa-up.sh — Clones The Whole Monorepo, Shells Out To openssl To Generate The Admission Webhook’s TLS Cert By Hand, Applies Raw Manifests
  2. charts/vertical-pod-autoscaler — An Actual Helm Chart Living In The Same Repo, Where Cert Generation Is A Helm Hook Job Instead Of A Shell Script Step You Run Yourself
git clone --depth 1 https://github.com/kubernetes/autoscaler.git
helm install vpa autoscaler/vertical-pod-autoscaler/charts/vertical-pod-autoscaler \
  --namespace vpa-system --create-namespace

Six Pods Come Up: Recommender x2, Updater x2, Admission-Controller x2 (Chart Default Replica Count, For HA). Combined Real-World Footprint After A Few Minutes Settled:

NAME                                    CPU(cores)   MEMORY(bytes)
vpa-...-admission-controller-cx99z      1m           10Mi
vpa-...-admission-controller-pjjlm      1m           10Mi
vpa-...-recommender-qsng2               2m           6Mi
vpa-...-recommender-zzgdv               1m           14Mi
vpa-...-updater-hn74p                   1m           13Mi
vpa-...-updater-kww9b                   2m           6Mi

~8m CPU, ~59Mi Memory, Total, Across All Six. On Nodes With 2 vCPUs Each, That’s Nothing. Whatever Hesitation I Had About Homelab Node Headroom Before Installing Turned Out To Be Unfounded.

Notes

  • Installing VPA Does Nothing By Itself. The Components Sit Idle Until You Create An Actual VerticalPodAutoscaler Object Pointing At A Workload. No Object, No Recommendation, No Resizing — Just Six Quiet Pods.

The Part I Didn’t Expect: The First Real Recommendation

Picked Headlamp (A Kubernetes Dashboard Deployment Already Running In The Cluster) As A Deliberately Low-Stakes Target — It Collects A Lot Of Cluster State Continuously, So It Seemed Like A Reasonable Thing To Watch. Attached A VPA Object In updateMode: "Off" — Recommendation-Only, Never Touches The Running Pod:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: headlamp-vpa
  namespace: headlamp
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: headlamp
  updatePolicy:
    updateMode: "Off"

Checked Actual Usage Right Before Applying It:

NAME                       CPU(cores)   MEMORY(bytes)
headlamp-fb548559f-zs8h2   2m           24Mi

Configured Resources On The Deployment At The Time: requests: 50m/64Mi, limits: 250m/256Mi. Already Comfortably Over-Provisioned Relative To Real Usage, By Eye.

Under A Minute After The VPA Object Landed:

status:
  recommendation:
    containerRecommendations:
      - containerName: headlamp
        lowerBound:   { cpu: 25m,  memory: 250Mi }
        target:       { cpu: 25m,  memory: 250Mi }
        upperBound:   { cpu: 640m, memory: "2013Mi" }

Both Targets Land At Roughly 10x Actual Usage: 25m Against 2m Of CPU, 250Mi Against 24Mi Of Memory. Only One Of Them Bothered Me. 250Mi, Sitting Six Mebibytes Under The Container’s Existing 256Mi Limit, Looked Like Something That Needed Explaining, And An Upper Bound North Of 2Gi For A Dashboard App Idling At 24Mi Made It Look Worse. 25m Looked Harmless, And Was Even Lower Than The Container’s Current 50m Request, So I Read It As A Sane Safety Margin And Moved On.

Noting That Deliberately, Because It Matters Later: I Went Digging Into The Memory Number And Never Checked The CPU One, For No Better Reason Than That It Looked Reasonable.

Why Memory Recommendations Are Built To Run High

Went Digging Into The Recommender’s Own Source Rather Than Guess:

  • Memory Targets Are Computed From The 90th Percentile Of Peak Usage Per Aggregation Window (Default 8 Days), Not Average Usage. CPU Is Compressible — Throttling A Container That Briefly Needs More CPU Just Slows It Down. Memory Isn’t — Exceed The Limit And The Kernel OOM-Kills The Container. The Recommender’s Math Is Deliberately Asymmetric For This Reason: It Would Rather Overestimate Memory Than Cause An OOM.
  • With Very Little Observation History, The Confidence-Interval Math Pulls The Bounds Apart On Purpose, Never Toward A Smaller Number. The Lower Bound’s Confidence Multiplier Collapses Toward The Target Itself When History Is Thin (Effectively: “Don’t Suggest Shrinking Something You Haven’t Watched Long Enough To Trust”). The Upper Bound’s Multiplier Does The Opposite — Balloons Toward An Enormous Number When History Is Thin, Which Is Exactly The ~2Gi Ceiling Showing Up Here. Both Directions Of “I’m Not Sure Yet” Get Resolved The Same Way: Toward Not Under-Provisioning.

What That Doesn’t Fully Explain, And What I’m Deliberately Not Papering Over: Why The Target Itself Landed At Almost Exactly 250Mi — Suspiciously Close To The Container’s Existing 256Mi Limit — Rather Than Somewhere Nearer The 24Mi It’s Actually Using Right Now. Didn’t Find A Clean, Citable Answer To That One In The Time I Spent On It. Possible It’s An Artifact Of How The Very First Recommendation Gets Computed Before A Full Aggregation Window Has Elapsed; Possible It’s Something Else Entirely. Leaving It Genuinely Open Rather Than Inventing A Tidy Explanation That Sounds Right But Isn’t Verified.

Notes

  • A Confident-Looking Number Isn’t The Same As A Confident Number. VPA Handed Back A Specific-Looking 250Mi, Not A Vague Range — But It Had Observed The Workload For Under A Minute. The Precision Of The Output Format Doesn’t Tell You Anything About How Much To Trust It Yet.
  • Memory And CPU Are Computed By Deliberately Different Maths. Memory Targets A High Percentile Of Peak Usage Specifically To Avoid OOMs; CPU Is Comparatively Relaxed Because Contention Only Throttles. Worth Knowing As Background — Though It Does Not, On Its Own, Explain The Specific Numbers Above.
  • The Number You Don’t Question Is The One To Watch. Both Targets Here Are About 10x Actual Usage, But Only The Memory One Looked Alarming Enough To Investigate. That Asymmetry Is In The Reader, Not In The Data.

What Happens Next

This Is Deliberately Posted Before The Story’s Finished. The Whole Point Of Confidence-Weighted Recommendations Is That They’re Supposed To Get Better With More History — VPA’s Own LowConfidence Condition Exists Specifically To Flag When A Recommendation Shouldn’t Be Trusted Yet. Letting This Run Against Real (Light) Traffic, Then Coming Back To See Whether The Memory Target Actually Converges Down Toward The 24Mi It’s Really Using, Or Whether It Stays Anchored Near The Configured Limit Regardless Of Real Usage. And While I’m There, Checking The CPU Number I Waved Through. Part 2 Covers Both.

Real-World Application

Scenario What To Do
Deciding Whether VPA Is Worth Installing Outside A Cloud-Managed Cluster It’s Fully Open-Source And Cloud-Agnostic — The Only Real Prerequisite Is Metrics Server. Check For It First; Some Distros (RKE2) Bundle It, Some Don’t
Installing VPA On Any Self-Managed Cluster Use The In-Repo Helm Chart Over vpa-up.sh Unless You Have A Specific Reason Not To — Same Components, Less Manual Cert Handling
Reading A Brand-New VPA Recommendation Check The RecommendationProvided / LowConfidence Conditions Before Acting On Any Number — A Recommendation Minutes Old Carries Very Different Weight Than One With Days Of History Behind It
A Recommendation That Looks Wrong Next To One That Looks Fine Check Both. Memory And CPU Are Computed By Different Maths (Peak Percentile Versus A More Relaxed CPU Path), But That Doesn’t Mean The Reasonable-Looking Number Has Been Verified — It Means It Hasn’t Been Questioned

Reference: