Situation

Had A Working GCP Weekly Digest Running On Google Apps Script For Weeks — Pulls Release Notes And Blog RSS Every Friday, Hands Them To Gemini To Write A Categorized Digest, Emails It Out, Archives A Copy To Google Drive. The Original Plan Was To Leave It Alone Until It Hit A Real Limit. Then A Different Reason To Migrate Showed Up: Wanted To Write This Up Next To The AWS Version As An “AWS vs GCP Serverless Digest Bot” Comparison, And Apps Script’s GmailApp/DriveApp/ScriptApp.getOAuthToken() Assume Every Reader Already Lives Inside A Google Account — Rebuilt It On Cloud Run Instead, Skipping The Original Migration Threshold On Purpose.

That Rebuild Alone Found Three Real Surprises And One Near-Miss. Then, With The Digest Itself Stable, Added The Other Half Of The Idea: An Opt-In Section That Points Google Cloud’s Own Recommender API At This Account And Has Gemini Turn The Findings Into Actual Advice — Deliberately Not Porting The AWS Sibling Project’s Hand-Written Detection Logic, Because GCP Already Ships A Better Version Of That Detection For Free. Getting It From “Written Against The Docs” To “Actually Tested” Surfaced Three More Problems, One Of Them The Kind That Looks Exactly Like Success Until You Know What To Check. And Finally, When SendGrid — The Email Provider This Whole Pipeline Depended On — Permanently Banned The Account, Replacing It With Gmail API Domain-Wide Delegation Found Two Last Blockers Neither Of Which Was A Bug In Any Line Of Code.

Result First:

Setup Cloud Run Job (Python 3.12) + Gemini 3.1 Flash-Lite (digest) + Cloud Recommender → Gemini (account advice) + Gmail API With Domain-Wide Delegation (delivery), Deployed From Source
Real Surprise 1 The SDK The Original Plan Assumed (google-cloud-aiplatform’s generative modules) Was Already Fully Retired — Not Deprecated-Soon, Actually Gone Two Months Before This Deploy
Real Surprise 2 The Model The V1 Version Used (gemini-2.5-flash-lite) Retires In Under Two Months — Copying It Forward Would’ve Broken A “Runs Every Friday” Job By Mid-October
Real Surprise 3 The GCP Blog RSS URL The Old Version Used Doesn’t Serve RSS Anymore — It 200s, But With An HTML Page, Not XML
Near-Miss Almost Ran A Local Gemini Auth Test Using Cached Credentials That Turned Out To Belong To A Completely Different Google Account And Project
Advice Bug 1 The Recommender API’s location wildcard "-" doesn’t work — a clean 400 InvalidArgument, not a permissions problem
Advice Bug 2 The first real verification run had every one of 42 (recommender × location) combinations fail with 403 (unset quota project + a disabled API) — and the script still reported “fetch succeeded, 0 results,” because a permission failure and a genuinely empty result look identical unless you specifically catch them differently
Advice Bug 3 The default location list was missing the one zone holding this project’s only VM and only disk — silently zero compute-related advice, ever, for that entire category
Mail Blocker 1 SendGrid permanently banned the account; org policy constraints/iam.disableServiceAccountKeyCreation then refused every attempt to download a Service Account key for the Gmail API replacement
Mail Blocker 2 With IAM, code, and Workspace delegation all correct, the first real send still came back 403 — Gmail API has not been used in project ... or it is disabled
Outcome All three pieces confirmed end-to-end on real infrastructure: digest generated and archived, account advice correctly flagging real findings, email delivered via Gmail API keyless delegation with a confirmed message ID

None Of The First Three Surprises Were Bugs In The Old Apps Script Code — They Were The Ground Having Moved Under A Ten-Week-Old Plan. The Advice Bugs And The Mail Blockers Are A Different Shape Entirely: Neither Came From Watching Something Break In Production. Both Came From Deliberately Verifying Assumptions Before Trusting Them — Because On This Particular Project, At This Particular Scale, Most Of These Failure Modes Would Otherwise Have Looked Exactly Like Success.

How It Broke (Root Cause)

Surprise 1: The SDK The Migration Plan Assumed No Longer Exists

The Original V2 Sketch Assumed google-cloud-aiplatform. Went To Actually Write The Call And Found:

google-cloud-aiplatform’s generative AI modules (vertexai.generative_models, etc.) were deprecated June 24, 2025, and fully removed June 24, 2026.

Two Months In The Past By The Time This Got Built. The Officially Recommended Path Is A Different Package Entirely — google-genai — Unifying Access To Both The Gemini Developer API And What Google Now Calls The “Gemini Enterprise Agent Platform” (Confirmed Independently Via gcloud services list --enabled, Which Lists aiplatform.googleapis.com Under The Title “Agent Platform API”, Not “Vertex AI API”). Then A Second, Smaller Trap: Two Web Searches Both Insisted vertexai=True “Does Not Appear Anywhere In The Document.” Installed The Package And Read It Directly Instead Of Trusting Either Search:

import inspect
from google import genai
print(inspect.signature(genai.Client.__init__))
# (self, *, enterprise=None, vertexai=None, api_key=None, credentials=None,
#  project=None, location=None, debug_config=None, http_options=None)

Both Parameters Exist. enterprise Is Current, vertexai Is Documented As A Legacy Alias. The Search Wasn’t Lying — It Just Answered With The New Name And Left Out That The Old One Still Quietly Works. Went With enterprise=True.

Surprise 2: The Model From The Old Version Retires In Under Two Months

gemini-2.5-flash-lite Is What The Apps Script Version Used, With Explicit Instructions To Keep Using It Unless There Was A Good Reason Not To. Checked Anyway:

Gemini 2.5 Flash-Lite shuts down October 16, 2026 (Gemini Developer API) / October 20, 2026 (Agent Platform API).

Seven Weeks Of Runway On A Job Meant To Run Indefinitely, Once A Week, Unattended. Switched To gemini-3.1-flash-lite, Currently Google’s Own “Most Cost-Efficient Gemini Model” Description — A Better Match For This Job’s Actual Usage Pattern Anyway.

Surprise 3: The Blog RSS Feed The Old Code Points At Just… Isn’t RSS Anymore

curl -sL -o /dev/null -w "%{http_code} %{content_type}\n" 'https://cloud.google.com/blog/rss/'
# 200 text/html; charset=utf-8

200, So Nothing In The Old Error Handling Would Have Flagged This — The Content Type Had Silently Swapped At Some Point After V1 Was Written. Found The Real Feed Instead Of Guessing At URL Variations:

curl -s 'https://cloudblog.withgoogle.com/rss/' | head -c 500
# <?xml version="1.0" encoding="utf-8"?>
# <rss version="2.0" xmlns:atom="..." xmlns:content="..." xmlns:dc="..." xmlns:media="...">

This Feed Carries The Same Namespaced Tags (dc:, content:, media:) That Broke The AWS Sibling Project’s Hand-Rolled Namespace-Stripping Regex. Didn’t Repeat That Bug. The Document Is Well-Formed XML With Its Namespaces Correctly Declared — ElementTree Handles That Natively. The Only Reason The AWS Version Needed String Surgery Was That It Manually (And Incompletely) Stripped The Declarations First. Skipped That Step Entirely; The Whole Bug Class Never Had A Chance To Exist Here.

Near-Miss: Testing Gemini Auth With The Wrong Google Account’s Cached Credentials

The Machine Already Had Application Default Credentials Cached From Unrelated Work On A Different Project. Checked First Instead Of Assuming:

TOKEN=$(gcloud auth application-default print-access-token)
curl -s "https://www.googleapis.com/oauth2/v1/tokeninfo?access_token=${TOKEN}"
# "email": "the-other-account@company.example"

Confirmed Wrong Account. Skipped The Local Test, Validated The Real Production Auth Path Instead — Through The Cloud Run Job’s Own Dedicated Service Account.

Building The Advice Half: Detection Belongs To Google, Not To Me

With The Digest Stable, Added An Opt-In Section: Point Google Cloud Recommender At This Project (Idle Static IPs, Idle Disks, Idle VMs, Machine-Type Rightsizing, Committed-Use Discounts, Near-Zero Project Utilization) And Have Gemini Turn The Findings Into Prioritized Advice. The AWS Sibling Project Had To Hand-Write String-Matching Detection Logic Against Cost Explorer Usage-Type Names, Because AWS Has No Free Equivalent (Trusted Advisor’s Cost Checks Need Business Support). GCP Ships That Detection For Free And More Completely Than Any Hand-Written Pattern List Would — So The Code Comment In account_context.py States The Principle Directly: detection belongs to whatever’s deterministic; the model only explains. On AWS, “Deterministic” Meant Hand-Written String Matching. On GCP, It’s Google’s Own API. Porting The AWS Approach Over Would Have Meant Rebuilding A Worse Version Of Something Google Already Gives Away.

Advice Bug 1: The Location Wildcard That Looks Documented But Isn’t

Most Recommenders Are Scoped Per Zone Or Per Region, And The Official API Reference Doesn’t Say Whether A Wildcard Location Works. Wrote The Code To Try - First And Fall Back To Enumerating Locations On Failure — Then Actually Tested It:

list(client.list_recommendations(
    parent=f'projects/{PROJECT}/locations/-/recommenders/{probe}'))
400 InvalidArgument: Invalid location: -

Clean Rejection, Not A Permissions Problem — So The Fallback Path Became The Only Path, And The Speculative Wildcard-First Branch Was Deleted Rather Than Left In As Dead Code Nobody Would Ever Trust Again.

Advice Bug 2: 42 Silent 403s That Reported Themselves As Success

First Real Verification Run Against The Actual Project. Every Single One Of 42 (Recommender × Location) Combinations Failed — An Unset Quota Project Plus A Not-Yet-Enabled Recommender API — And The Original Code Caught All Exceptions The Same Way, So The Result Was A Clean Empty List. The Script’s Own Log Line Read:

[advice] 掃過 6 個 recommender × 9 個 location,42 個組合不適用,取得 0 筆建議

“0 Results” And “42 Silent Permission Failures” Produce The Exact Same Output Shape If Every Exception Gets Caught The Same Way. That’s The Trap: A Genuinely Empty Recommendation List Is A Real, Expected, Even Good Outcome For A Personal Project — So The Code Cannot Simply Alarm On Empty. It Has To Tell The Difference Between “Nothing To Report” And “Never Actually Asked” Without A Human Reading Every Response. The Fix Splits Exceptions Into Two Classes:

_SKIPPABLE = (gexc.InvalidArgument, gexc.NotFound)      # this combination doesn't apply — expected
_FATAL     = (gexc.PermissionDenied, gexc.Unauthenticated, gexc.ResourceExhausted)  # environment is broken — must surface

Plus A Last-Resort Guard: If Every Combination Gets Skipped, That’s Not “No Recommendations” — It’s “Never Successfully Asked Once,” Almost Certainly A Misconfigured Location List, And It Now Raises Instead Of Returning An Empty List That Would Have Rendered As “No Open Recommendations This Week” — Indistinguishable From The Genuinely Good Outcome It’s Impersonating.

Advice Bug 3: The Location List That Silently Excluded The Only Resources That Existed

Fixing Bug 2 Meant The Fetch Actually Succeeded — And Then Returned Zero Compute-Related Recommendations Anyway, For A Different Reason. The Default Location List Didn’t Include The One Zone Holding This Project’s Only VM And Only Disk. Every Compute-Related Recommender Query Landed On Locations With Nothing In Them — Technically Successful, Structurally Guaranteed To Find Nothing. A Location List Missing One Entry Produces Exactly The Same Output As A Project With No Idle Resources. There’s No Error For “You Forgot A Zone” — Only The Absence Of Findings That Should Have Existed. Caught Only By Cross-Checking The List Against Where This Project’s Actual Resources Live, Not By Anything The API Itself Could Have Signaled.

Mail Blocker 1: SendGrid Banned The Account, Then The Org Wouldn’t Let Anyone Download A Key

SendGrid Permanently Banned The Account — Two Delivery Deferrals At Exactly 72.01 Hours Each, Followed By A Reactivation Request Denied For Good. Since kmp.tw’s MX Records Already Point At Google, Gmail API With Domain-Wide Delegation Looked Like The Natural Replacement. Tried The Standard Recipe First — Dedicated Mailer Service Account, Domain-Wide Delegation Enabled, Download The JSON Key, Mount Through Secret Manager. gcloud iam service-accounts keys create Refused Outright:

“If the iam.disableServiceAccountKeyCreation constraint is enforced for your organization, then you can’t create keys for any service accounts in your organization.”

Not This Service Account — Any Service Account, Org-Wide. The Keyless Alternative Uses The Cloud Run Job’s Own Runtime Identity, Granted roles/iam.serviceAccountTokenCreator Scoped Only To The Mailer SA Resource (Not Project-Wide). At Send Time, The IAM Credentials API’s signJwt Signs An RFC 7523 JWT-Bearer Assertion (Subject wei@kmp.tw — This Is The Actual Delegation Step) Which Gets Exchanged At oauth2.googleapis.com/token For A Genuine, Short-Lived Access Token. No Private Key Is Ever Generated, Downloaded, Or Written To Disk Anywhere In The Chain.

Mail Blocker 2: 403 After Every Other Piece Was Already Right

Once Workspace Admin Console Authorization Was Done, The Token Exchange — Which Had Been Failing With 401 unauthorized_client The Whole Time Delegation Wasn’t Authorized — Started Returning Real Tokens. The Very Next Call, users.messages.send, Came Back 403:

Gmail API has not been used in project 354846332057 before or it is disabled.
Enable it by visiting
https://console.developers.google.com/apis/api/gmail.googleapis.com/overview?project=354846332057
then retry. If you enabled this API recently, wait a few minutes for the
action to propagate to our systems and retry.

The Entire Identity Chain Was Already Working. The Only Thing Wrong Was That gmail.googleapis.com Had Simply Never Been Enabled On This Project Before, Because Nothing In It Had Ever Called Gmail API Until This Line Of Code:

gcloud services enable gmail.googleapis.com --project=amp-project-161400

Reran The Job. Real Send, Real Log Line:

[Gmail] 已送出,message id:1a067facf2c866b7
Email 已寄出
完成!

Worth Noting: The Project Identifier In That URL (354846332057) Is The Project Number, Not The Project ID (amp-project-161400) — Searching Logs For The Project ID String Won’t Find This Error At All.

Notes

  • A rebrand can be real even when it sounds like a hallucination. aiplatform.googleapis.com Titled “Agent Platform API,” Vertex AI’s Own Docs Saying They’re “No Longer Being Updated” — Both Confirmed Independently Through gcloud services list, Not Assumed From A Search Summary.
  • When a web search and the actual installed package disagree, install the package. Ten Seconds Of inspect.signature() Settled What Two Searches Got Confidently Wrong.
  • “It returns 200” is not the same as “it still works.” True Of A Dead RSS Feed Silently Serving HTML, And True Of An API Call That Reports “Success” While Having Caught Every Failure The Same Way.
  • The scariest bugs in this project all share one shape: a genuinely good outcome and a broken one produce identical output. An Empty Recommendation List Because There’s Nothing To Flag, And An Empty List Because Every Call 403’d — Both Print The Same Sentence Unless The Code Is Specifically Built To Tell Them Apart. A Missing Location In A Config List Produces The Same Silence As A Project With No Idle Resources.
  • Domain-wide delegation fails at two independent points that don’t look alike. Missing Workspace Authorization Fails At Token Exchange With 401 unauthorized_client. An API Never Enabled On The GCP Project Fails Later, At The Actual Call, With 403 SERVICE_DISABLED. Fixed In Two Completely Different Consoles By Two Completely Different People.
  • Porting a sibling project’s workaround without asking why it existed there imports the wrong lesson. The AWS Version Needed Hand-Written Detection Because AWS Has No Free Equivalent. Copying That Pattern To GCP Would Have Meant Reinventing A Worse Version Of Something Google Already Provides.

How To Prevent It

Scenario What To Do
Migrating code that calls a managed AI SDK, based on a plan written more than a few weeks ago Re-check the SDK’s current deprecation status before writing a single line
A model ID was chosen previously with instructions to “keep using it unless there’s a reason not to” Check the model’s specific retirement date, not just whether it works today
An RSS/feed URL worked when a project doc was first written Re-verify with curl -sL -o /dev/null -w "%{http_code} %{content_type}" before reusing it
Porting a workaround from a sibling project on a different cloud Check whether the reason it existed still applies — sometimes the platform already solved it natively
A machine has cached credentials for multiple accounts across unrelated projects Verify identity (tokeninfo, etc.) before the first real API call in a new session
Catching every exception from a bulk API scan the same way Split “this combination doesn’t apply” from “the environment is broken” — an empty result and a wall of silent permission failures must not look identical
A config list enumerates locations/regions/zones to scan Cross-check it against where resources actually live — a missing entry produces silence, not an error
An org enforces constraints/iam.disableServiceAccountKeyCreation Don’t request an exception — go keyless via IAM Credentials API signJwt with a narrowly-scoped serviceAccountTokenCreator grant
A Google API call fails with SERVICE_DISABLED / “has not been used in project” Check gcloud services list --enabled before assuming it’s another auth problem

Reference: