Situation

Had Already Set Up VPA On A GKE Workload (A ClickHouse Backend, After It OOM’d Twice) And Assumed — Wrongly — That Vertical Pod Autoscaler Was Something GKE Provided As A Managed Feature. Turns Out The “Enable” Button On GKE Is Just Google’s Wrapper Around The Same Open-Source kubernetes/autoscaler Project Anyone Can Install Anywhere. Wanted To Confirm That For Real, So I Put It On A 6-Node Homelab RKE2 Cluster With No Cloud Provider In Sight.

It Went In Clean. What Wasn’t Clean Was The First Recommendation It Handed Back For A Live Workload — Ten Times The Container’s Actual Memory Usage, Within Minutes Of Being Installed.

Continue reading

Situation

Had Been Watching A LinkedIn Account Post AWS Updates Every Week For A While — Always Just Updates, And Kept Wondering If There Was A More Human Way To Deliver Them. So, Riding The AI Wave, Built A Lambda + Bedrock + SES + S3 Pipeline: Every Friday It Pulls The AWS What’s New And Blog RSS Feeds, Hands Them To Amazon Nova Pro To Write A Categorized Digest, Emails It Out, Archives A Copy To S3 — And Also Has Claude Opus Look At This Account’s Actual Resource Usage And Suggest Optimizations.

Continue reading

Situation

Needed To Wire All Of Akamai’s Data Into An Existing Grafana On GKE. Split Into Two Dashboards: CDN Data Rode On A Plugin That Was Already There, And WAF/Bot Logs Went Through A Small Proxy I Wrote Myself, Because The Only Grafana Plugin That Speaks Akamai’s SIEM API Is A Paid Third-Party One. Didn’t Want To Pay For It. (Cheapskate.)

Both pieces hit the same kind of bug: something that looks fixed, deploys clean, and still doesn’t work — with nothing useful in the logs to explain why.

Continue reading

Situation

After I got Keycloak wired up, logging into the dashboard started getting stuck. No error, no crash log. Just one of those situations where you feel like you’re missing something obvious and can’t see it. Today I tried again, and this time there was a clear error in the log: Invalid parameter: redirect_uri

Continue reading

Situation

A Recent Task Was To Convert An Existing External Environment On GCP Into An Internal One. The Whole Setup Runs On GKE With Gateway API Handling Load Balancing. The Assumption Going In Was That A Google-Managed Certificate Would Attach To The Gateway API Internal LB The Same Way It Already Worked On The Gateway API External LB. That Assumption Ran Into Trouble Immediately.

Continue reading

Situation

My Home Lab Originally Only Used LLDAP For User Authentication, But Since The Systems I Wanted To Integrate Later Only Spoke OIDC, I Put Keycloak In Front Of It As A Bridge, With Everything Else Talking To Keycloak Instead. Wiring That Up And Testing It Turned Up Something Interesting: Got The LDAP Federation Wired Up, Got The OIDC Client Configured, Got Kubernetes Itself Trusting The Issuer. Then, Testing Whether Any Of It Actually Enforced Anything, I Ran One kubectl Command With A Made-Up Token String, And Got Back A Full Pod List. For About Thirty Seconds I Was Convinced I’d Found A Cluster-Wide Auth Bypass.

I Hadn’t. The Bug Was In My Test, Not My Cluster.

Continue reading

Situation

My Home RKE2 Cluster Had Envoy Gateway Handling North-South Traffic, Routed To Whatever Service Needed It. It Worked Fine. What It Didn’t Have Was Any Answer For East-West: The Handful Of Services Calling Each Other Internally Were Doing It In Plaintext, No Identity, No mTLS. I’d Been Meaning To Add A Service Mesh For A While, Mainly Because This Is Exactly The Kind Of Thing A Real Company Would Actually Do, And “Migrate A Live Ingress From One Implementation To Another, Zero Downtime” Is About As Real-Company As It Gets. Decided To Fully Replace Envoy Gateway With Istio’s Own Gateway In Ambient Mode, Rather Than Just Bolting Istio’s Mesh On The Side, And Did It The Same Day.

Continue reading

Author's picture

Gordon wei

Stay Hungry Stay Foolish

iKala Cloud Solution Engineer | AWS Community Builders

Taiwan