[K8S] Migrating From Envoy Gateway To Istio Ambient
Situation
My Home RKE2 Cluster Had Envoy Gateway Handling North-South Traffic, Routed To Whatever Service Needed It. It Worked Fine. What It Didn’t Have Was Any Answer For East-West: The Handful Of Services Calling Each Other Internally Were Doing It In Plaintext, No Identity, No mTLS. I’d Been Meaning To Add A Service Mesh For A While, Mainly Because This Is Exactly The Kind Of Thing A Real Company Would Actually Do, And “Migrate A Live Ingress From One Implementation To Another, Zero Downtime” Is About As Real-Company As It Gets. Decided To Fully Replace Envoy Gateway With Istio’s Own Gateway In Ambient Mode, Rather Than Just Bolting Istio’s Mesh On The Side, And Did It The Same Day.
Result First:
| Before | Envoy Gateway, north-south only, no mesh |
| Decision | Fully replace with Istio Gateway (not “add mesh, leave ingress alone”) — same mTLS either way, this route also removes a system and doubles as ingress-migration practice |
| Services Migrated | 4 — Gitea (plain manifest), Bifrost/Headlamp/Keycloak (GitOps-managed) |
| Downtime | Zero |
| Biggest Surprise | The existing wildcard TLS cert needed zero changes — certificateRefs is a Gateway API field, not an Envoy or Istio proprietary mechanism |
| Actual Risk | Not the TLS cert (that was fine) — it was almost editing GitOps-managed routes by hand and getting silently reverted |
The Two Routes Look Different In Effort, Not In What You Get
Without Comparing Service Mesh Options Against An Existing Gateway API Setup, It’s Easy To Assume Ripping Out The Current Ingress Controller Is What “Unlocks” Mesh Features Like mTLS. Gateway API (Who Handles Ingress) And Service Mesh (Who Handles Service-To-Service) Are Separate Concerns. You Can Install Istio’s Data Plane Next To An Untouched Envoy Gateway And Get Full mTLS Between Meshed Pods, No Ingress Changes Required. The Only Reason To Additionally Swap The Gateway Itself Is Consolidation (One Control Plane Instead Of Two) And, In This Case, The Migration Practice. If Someone’s Evaluating This For A Real Team And The Answer They’re Looking For Is “Which Gives Us More Mesh Capability” — Neither, They’re The Same. The Question Is Really “Do We Also Want To Consolidate Ingress.”
The Data Plane, Meanwhile, Went With Ambient Mode Over Sidecars. A Node-Level ztunnel Handles mTLS/L4 Instead Of A Proxy Container In Every Pod. My Nodes Aren’t Great, And Ambient’s Whole Pitch Is Lower Overhead At That Scale.
Trap One (That Wasn’t): The Wildcard Cert Just Worked
I Expected To Re-Issue Or At Least Re-Mount The Existing *.internal Wildcard TLS Cert When Switching Gateway Implementations. It Turns Out certificateRefs — The Field That Points A Gateway’s TLS Listener At A Kubernetes Secret — Is Part Of The Gateway API Spec Itself, Not Something Envoy Gateway Or Istio Each Bolt On Differently. Same Secret, Same Field Name, Different Controller Reading It:
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: istio-gateway
namespace: istio-system
spec:
gatewayClassName: istio
listeners:
- name: https
port: 443
protocol: HTTPS
tls:
mode: Terminate
certificateRefs:
- kind: Secret
name: wildcard-tls # same Secret the old Envoy Gateway used
namespace: envoy-gateway-system
The Only Thing Needed Was A ReferenceGrant In The Secret’s Namespace, Because The New Gateway Object Lives In A Different Namespace (istio-system) Than Where The Cert Secret Happened To Be Parked:
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata:
name: allow-istio-system-read-tls
namespace: envoy-gateway-system
spec:
from:
- group: gateway.networking.k8s.io
kind: Gateway
namespace: istio-system
to:
- group: ""
kind: Secret
name: wildcard-tls
Worth Calling Out: I’d Assumed Before Starting That A ReferenceGrant Was Already In Use Somewhere In The Existing Setup. It Wasn’t. The Old Gateway And The Secret Had Always Lived In The Same Namespace, So Cross-Namespace Referencing Had Simply Never Come Up Before. Good Reminder That “This Pattern Is Probably Already Handled Somewhere” Is A Guess, Not A Fact, Until You Actually Grep For It.
Trap Two: GitOps-Managed Routes Will Silently Revert Your Manual Edit
Three Of My Four Services (An AI Gateway, A K8s Dashboard, And An Identity Provider, I’ll Call Them Bifrost, Headlamp, And Keycloak, Their Real Names) Had Their HTTPRoute Managed By Flux — Which I Hadn’t Front-Of-Mind When I Started Testing. The Fourth, A Self-Hosted Gitea, Was A Plain Manual Manifest, So I Just kubectl patch-Ed It Directly, No Problem. For The Other Three, Editing The Live Object By Hand Would Have Worked For Exactly As Long As It Took Flux’s Next Reconcile Loop To Notice The Drift And Put It Back:
# This gets reverted on the next reconcile if the Kustomization has prune: true
# and the source repo still has the old parentRef
spec:
parentRefs:
- name: main-gateway # <- edit this live, Flux puts it right back
namespace: envoy-gateway-system
The Actual Fix Was Editing The Source Repo, Not The Cluster:
git clone <gitops-repo>
# change parentRefs from main-gateway to istio-gateway in the HTTPRoute manifest
git commit -m "Migrate HTTPRoute to Istio Gateway"
git push
kubectl annotate kustomization <name> -n flux-system \
reconcile.fluxcd.io/requestedAt="$(date +%s)" --overwrite
None Of This Is Istio-Specific. It’s The Standard “Where’s The Actual Source Of Truth” Question A GitOps Setup Forces On Every Infra Change. But It’s Easy To Forget In The Moment When You’re Focused On Testing A New Gateway Config And Just Want To See If It Works.
Validating Without Touching DNS: Run Both Gateways In Parallel
The Zero-Downtime Part Came From Never Touching Production Routing Until The New Path Was Already Proven. MetalLB Handed The New Istio Gateway Its Own IP, Separate From The Old One, So Both Gateways Ran Side By Side. Testing The New Path Didn’t Require Touching DNS At All — curl --resolve Fakes The DNS Answer For A Single Request:
# hits the NEW gateway directly, real hostname for correct SNI/cert matching
curl --cacert ca.pem --resolve myservice.internal:443:<new-gateway-ip> \
https://myservice.internal/
# 200
# side-by-side against the OLD gateway, still serving real production traffic
curl --cacert ca.pem --resolve myservice.internal:443:<old-gateway-ip> \
https://myservice.internal/
# 200
Only After Every Service Came Back 200 On The New Path Did I Touch The Authoritative DNS Records (Two Servers, Both Had To Be Updated And Both Individually Verified — More On Why Below) And Point Them At The New Gateway’s IP. Repeated Per Service: Confirm On The New Path → Flip DNS For That One Service → Confirm Again On The Real Hostname → Move To The Next. Only After All Four Were Confirmed Live On The New Gateway Did I Uninstall The Old One.
Notes
certificateRefsIs A Gateway API Spec Field, Not An Envoy Or Istio Proprietary Mechanism. Switching Gateway Implementations Doesn’t Require Touching TLS Certs At All, As Long As The New Controller Also Implements Gateway API (Most Do). Confirm This Before Assuming A Cert Migration Task Even Exists.- A
ReferenceGrantYou’ve Never Needed Before Doesn’t Mean It’s Already Handled Elsewhere. It Might Just Mean The Situation Requiring It Never Came Up. Check, Don’t Assume. - GitOps-Managed Resources Revert Manual Edits On The Next Reconcile. This Applies To Any Ingress/Route Migration In A Flux Or ArgoCD-Managed Cluster, Not Just This One — Identify Which Resources Are GitOps-Owned Before Touching Anything Live.
curl --resolveLets You Fully Validate A New Ingress Path — TLS, Routing, Backend — Without Touching DNS Or Risking Live Traffic. Flip DNS Only After Every Service Independently Passes On The New Path.- Actual Resource Overhead Was Smaller Than Expected. My Cluster’s Nodes Are Tight On Memory (One Physical Host Running All Five VMs, Already Low On Free RAM). I Was Ready To Abort Or Scale Back If Ambient’s Per-Node
ztunnelMade That Worse. Real Numbers: 1-4 Percentage Points Of Memory Increase Per Node. Worth Measuring Before Assuming A Resource-Constrained Environment Can’t Handle It.
Real-World Application
| Scenario | What To Do |
|---|---|
| Evaluating Whether To Add A Service Mesh On Top Of An Existing Gateway API Setup | Remember Gateway API And Service Mesh Are Orthogonal — You Can Add Mesh Capability (mTLS, Tracing) Without Touching Ingress At All. Only Consolidate The Two If You Specifically Want One Fewer System To Run |
| Migrating Ingress In Any Cluster Using Flux Or ArgoCD | Identify Which Routes Are GitOps-Managed Before Touching Anything. Edit The Source Repo, Not The Live Object — Manual Edits To Managed Resources Get Silently Reverted On The Next Reconcile |
| Any Zero-Downtime Ingress/LB Migration (Not Just Gateway API) | Run Old And New In Parallel On Separate IPs, Validate The New Path With curl --resolve Or Equivalent Per-Request DNS Override, And Only Flip Real DNS After Every Route Independently Passes |
| Resource-Constrained Clusters Considering Ambient Mode | Don’t Assume The Overhead Is Prohibitive Without Measuring — Take A kubectl top nodes Baseline Before Installing, Compare After. The Actual Number May Be Smaller Than The Sidecar-Model Reputation Suggests |
Reference:
- Istio — Ambient Mode
- Istio — Gateway API Support
- Real Incident: Home Lab RKE2 Cluster, Envoy Gateway → Istio Ambient Migration