Situation

Had Been Watching A LinkedIn Account Post AWS Updates Every Week For A While — Always Just Updates, And Kept Wondering If There Was A More Human Way To Deliver Them. So, Riding The AI Wave, Built A Lambda + Bedrock + SES + S3 Pipeline: Every Friday It Pulls The AWS What’s New And Blog RSS Feeds, Hands Them To Amazon Nova Pro To Write A Categorized Digest, Emails It Out, Archives A Copy To S3 — And Also Has Claude Opus Look At This Account’s Actual Resource Usage And Suggest Optimizations.

sam build Passed. Template Validated. Everything Looked Done On Paper. First Real Deploy Still Found Two Bugs That Zero Amount Of Re-Reading The Code Would Have Surfaced, Because Both Only Show Up Once Real AWS Resources Are Talking To Each Other. The Account-Advice Feature Had A Longer Road: Two Prompt Versions That Produced Nothing Useful, A Model Swap That Revealed A Genuine Quality Gap Once The Prompt Was Actually Good, And Two More Bugs That Only Surfaced When Prepping To Open-Source The Project And Asking Whether Any Of It Would Hold Up On A Bigger Account.

Result First:

Setup Lambda (Python 3.12, arm64) + Amazon Bedrock (Nova Pro For The Digest Body, Claude Opus 5 For Account Advice, Both Via Cross-Region Inference) + SES + S3, Deployed Via SAM
First Deploy Blocker CAPABILITY_IAM Wasn’t Enough — Template Declares An Explicit IAM RoleName, Which Bumps The Requirement To CAPABILITY_NAMED_IAM
Bug 1 IAM Policy For bedrock:InvokeModel Was Scoped To The Wrong ARN — Substituted The Cross-Region Inference Profile ID Straight Into The foundation-model Resource, Which Isn’t How The Underlying Model ARN Is Actually Shaped
Bug 2 RSS Parser Silently Returned Zero Blog Posts Every Single Run — A Namespace-Stripping Regex Only Half Did Its Job
Two Failed Advice Prompts Service names alone produced five lines of pure filler; adding usage-type granularity gave real signal, but the model still missed the two most valuable findings
Model Comparison Same already-good prompt, same data, only the model swapped — Nova Pro forced a “savings” framing on 4 of 5 findings despite an explicit instruction not to; Claude Opus 5 cross-referenced multiple signals into real findings and correctly stayed quiet where there was nothing to save
Bug 3 GetCostAndUsage caps at 1,000 groups per response with no pagination handling — a bigger account would silently lose entries with zero error
Bug 4 The pattern-flagging step scanned only the already-truncated top-4-per-service display list — the smallest, most-signal-bearing usage types were exactly the ones most likely to fall outside it
How I Found Them The first two by actually invoking the deployed function and reading CloudWatch Logs line by line, not just checking StatusCode: 200. The last two by asking “would this hold up on a bigger account?” before shipping the code to anyone else’s AWS account — not by watching anything break
Outcome End-to-end run confirmed: 54 What’s New items + 3 blog posts parsed, Nova Pro digest generated, Opus 5 account advice generated (flag count on the same account went 7 → 14), saved to S3, emailed via SES

Neither Of The First Two Bugs Broke sam build. Neither Broke sam validate. Both Only Existed The Moment Real IAM Policies Met Real Bedrock APIs And Real RSS Bytes. The Last Two Were Quieter Still — This Account Is Structurally Too Small To Ever Trigger AWS’s Own Pagination Cap, So Both Were Found By Asking, Not By Watching Something Fail.

How It Broke (Root Cause)

Bug 1: The Cross-Region Inference Profile ARN Trap

Nova Pro On Bedrock Is Set Up To Use Cross-Region Inference — You Call It With An Inference Profile ID Like us.amazon.nova-pro-v1:0 Instead Of A Plain Model ID, And Bedrock Load-Balances The Actual Request Across A Handful Of Regions Behind The Scenes. The IAM Policy I’d Written Assumed The Profile ID Was Also The Model’s Identity For Permission Purposes, So It Built The Resource ARN Like This:

Resource:
  - !Sub 'arn:aws:bedrock:${BedrockRegion}::foundation-model/${BedrockModelId}'
  # BedrockModelId = us.amazon.nova-pro-v1:0

First Real Invoke After Deploy:

AccessDeniedException: User: arn:aws:sts::...:assumed-role/aws-weekly-digest-role.../aws-weekly-digest
is not authorized to perform: bedrock:InvokeModel on resource:
arn:aws:bedrock:us-east-1::foundation-model/amazon.nova-pro-v1:0

Look Closely At That Resource — No us. Prefix. When You Invoke Through A Cross-Region Inference Profile, IAM Doesn’t Just Check Permission On The Profile Itself. It Also Checks Permission On The Underlying Foundation Model That The Profile Actually Routes The Request To — And That Underlying Model’s ARN Never Carries The us./eu./apac. Prefix. That Prefix Only Exists On The Profile ID, Not On The Real Model.

Didn’t Want To Guess At Which Regions It Actually Routes To, So Asked The API Directly Instead Of Assuming:

import boto3
c = boto3.client('bedrock', region_name='us-east-1')
r = c.get_inference_profile(inferenceProfileIdentifier='us.amazon.nova-pro-v1:0')
for m in r['models']:
    print(m['modelArn'])
arn:aws:bedrock:us-east-1::foundation-model/amazon.nova-pro-v1:0
arn:aws:bedrock:us-west-2::foundation-model/amazon.nova-pro-v1:0
arn:aws:bedrock:us-east-2::foundation-model/amazon.nova-pro-v1:0

Three Destination Regions, Not One. The Original Policy Only Granted us-east-1, So Even Fixing The Model ID Without Fixing The Region Coverage Would Have Left An Intermittent Failure Waiting To Happen Depending On Which Region Bedrock Decided To Route To That Day. Fixed It By Explicitly Granting The Real Foundation-Model ARNs With A Region Wildcard, Plus The Inference-Profile ARN Separately:

Resource:
  - 'arn:aws:bedrock:*::foundation-model/amazon.nova-pro-v1:0'
  - 'arn:aws:bedrock:*::foundation-model/amazon.nova-lite-v1:0'
  - 'arn:aws:bedrock:*::foundation-model/anthropic.claude-3-5-sonnet-20241022-v2:0'
  - 'arn:aws:bedrock:*::foundation-model/anthropic.claude-3-haiku-20240307-v1:0'
  - !Sub 'arn:aws:bedrock:${BedrockRegion}:${AWS::AccountId}:inference-profile/*'

Redeployed, Reinvoked — Clean Pass Through Bedrock On The Second Try.

Bug 2: The RSS Parser That Silently Gave Up On Every Blog Post

The AWS Blog RSS Feed Declares A Pile Of XML Namespaces On Its Root <rss> Tag — dc:, content:, slash:, atom:, And So On. To Dodge Dealing With Namespace-Aware ElementTree Parsing, The Original Code Just Regex-Stripped The xmlns Declarations Off The Root Element Before Parsing:

clean = re.sub(r'\s+xmlns(?::[a-z]+)?="[^"]*"', '', raw[start:])
root = ET.fromstring(clean)

Problem Is, Stripping The Declarations Doesn’t Strip The Prefixed Tags Still Sitting Inside Every <item> — Things Like <dc:creator>Jeff Barr</dc:creator> Are Still There. Once The xmlns:dc="..." Declaration That Would Have Told The Parser What dc: Means Is Gone, ElementTree Hits The First <dc:...> Tag And Throws:

unbound prefix: line 5, column 1

The Function Already Had A Try/Except Around The Whole Fetch, By Design — If The Blog Feed Is Down Or Malformed, The Digest Should Still Ship With Just The What’s New Section Instead Of Failing The Whole Run:

except Exception as e:
    print(f'[Blog] 跳過:{e}')
    return []

Which Is Exactly What It Did. Every Single Invoke, Silently, From The First Line Of Code. statusCode: 200, Email Sent, S3 File Written — Just With A Permanently Empty Blog Section, And No Symptom Visible Anywhere Except A CloudWatch Log Line You’d Only See If You Went Looking.

Fix Wasn’t Complicated Once Found — Strip The Prefix Off The Tags Too, Not Just The Declarations:

clean = re.sub(r'\s+xmlns(?::[a-z]+)?="[^"]*"', '', raw[start:])
clean = re.sub(r'</?[a-zA-Z0-9]+:', lambda m: m.group(0).replace(':', ''), clean)
root = ET.fromstring(clean)

<dc:creator> Becomes <creator>, </dc:creator> Becomes </creator>, And Since The Code Only Ever Reads title / link / description / pubDate Off Each Item Anyway, Renaming Tags It Never Touches Costs Nothing. Redeployed, Reinvoked — Blog Count Went From 0 篇 To 3 篇 In The Log Line.

Two Failed Prompts Before Account Advice Was Worth Reading

Once The Digest Itself Was Stable, Added An Opt-In Section: Feed This Account’s Actual Cost Explorer Usage To Bedrock, Append Optimization Advice To The Email. Version 1 Fed The Model Nothing But Service Names And Got Back Five Lines Of Pure Filler:

[AWS WAF]:現況推測:可能沒有設定 WAF 規則來保護應用程式。建議動作:設定 WAF 規則⋯⋯

“Possibly Not Configured, Recommend Configuring” — True Of Any Account, Because The Model Had Nothing But A Name To Guess From. Adding A Second GroupBy Dimension On USAGE_TYPE Changed The Input Entirely:

Amazon Virtual Private Cloud
    APS1-PublicIPv4:IdleAddress = 0.0892      ← the name says Idle
Amazon Elastic Container Service for Kubernetes
    APE2-AmazonEKS-Hours:extendedSupport = 0.479   ← paying extra to run an EOL version

Real Signal Now — But The Model Still Missed The Two Most Valuable Findings And Spent A Slot On A Low-Value CloudFront Check That Wasn’t Backed By Anything In The Data. Fix Moved Detection Out Of The Prompt Entirely: String-Match Known-Waste Usage-Type Patterns (IdleAddress, extendedSupport, NatGateway-Hours) In Code, Flag Them Before The Model Ever Sees Them, Let The Model Only Explain What’s Already Flagged. Third Version Got All Five Right.

Nova Pro vs Claude Opus 5: Same Data, Two Different Levels Of Reasoning

With The Prompt Already At Its Best Version, Swapping Only BEDROCK_MODEL_ID From Nova Pro To Claude Opus 5 (us.anthropic.claude-opus-5) Produced Two Different Kinds Of Output From The Same Data.

Nova Pro: Four Of Five Findings Led With “Expected Effect: Cost Savings” — Directly Against What The Prompt Said — On An Account That Spent $0.0014 Over Ninety Days. Each Finding Explained In Isolation, No Cross-Referencing.

Claude Opus 5: 2,160 = 24 × 90 → Two Regions Each Running A DynamoDB Table Continuously For The Full Window. extendedSupport Hours Exactly Equal To Total Runtime → The Cluster Launched On An Outdated Version From Day One, Not Aged Into One. Zero BoxUsage In The Regions Holding The Flagged Snapshots → The Instances Are Long Gone, Only Backups Are Billing. And Correctly Stayed Quiet Once IPv4 Usage Matched The Load Balancer Count.

Cost, Measured, Not Estimated:

Model Per Run Per Year (52 Runs)
Nova Pro $0.0035 $0.18
Claude Opus 5 $0.164 (in 4,296 / out 5,702, thinking included) $8.53

47x More Expensive, $8.35 More A Year In Absolute Terms — Against The Cost Explorer API Call Itself Running $0.01/Week ($0.52/Year). At This Job’s Scale, Model Choice Was Never A Cost Problem, Only A Quality One. Production Switched To Opus 5 After This Comparison.

Switching Wasn’t Free, Though. Two Things Broke On The First Real Call: The Existing Claude Branch Sent A temperature Parameter On Every Call, And Opus 5 Has Dropped Sampling Parameters Entirely — A Flat 400. The Bare Model ID Also 400’d (“On-demand throughput isn’t supported”) Until Switched To The Same Inference-Profile Form Nova Already Used (us.anthropic.claude-opus-5). Adaptive Thinking Pushed A Single Advice Call To 82.9 Seconds, Which — Stacked On Top Of The Digest’s Own ~118 Seconds — Meant The Lambda Timeout Had To Double From 300s To 600s.

Two Bugs A $0.0014 Account Was Never Going To Trigger

Prepping To Open-Source The Project Meant Asking A Question The Account Running It Couldn’t Answer On Its Own: What Happens At Actual Scale?

Bug 3 — No Pagination On Cost Explorer Results. GetCostAndUsage Caps At 1,000 Groups Per Response, The Original Code Only Ever Read The First Page. The Failure Is The Dangerous Kind: The List Comes Back Shorter, With Nothing Indicating Anything Was Cut. Fixed With A Straightforward Pagination Loop.

Bug 4 — Flagging Ran Against The Already-Truncated Display List. The Display Step Trims Each Service To Its Top 4 Usage Types To Keep The Prompt Readable — But The Flagging Function Was Iterating That Same Trimmed List, Meaning The Smallest, Most Valuable Signals Were Exactly The Ones Most Likely To Be Cut First (An Idle IP At Usage 0.089 Sits Nowhere Near A Top-4-By-Cost Cutoff). Fixed By Scanning The Full List For Flags While Still Only Displaying The Top 4 Per Service. Flag Count On The Same Test Account Went From 7 To 14.

Notes

  • A gracefully-degrading except block is also a great place for a bug to live forever, invisibly. The whole point of catching that exception was so one broken feed wouldn’t take down the whole digest — which worked exactly as designed. But “worked as designed” and “worked correctly” turned out to be two different things, and only reading the actual log line told them apart. StatusCode: 200 was never going to say that.
  • IAM permission errors for Bedrock cross-region inference profiles don’t show up until the first real InvokeModel call. sam validate doesn’t know what ARN a profile ID actually resolves to at runtime — that information only exists behind the GetInferenceProfile API, so there’s no way to catch this at template-authoring time short of actually calling it and checking.
  • CAPABILITY_IAM vs CAPABILITY_NAMED_IAM is a one-line fix but a fully opaque error the first time you hit it. The moment a SAM template gives an AWS::IAM::Role an explicit RoleName (instead of letting CloudFormation generate one), the required capability silently changes.
  • A model producing generic advice is a signal problem before it’s a model problem. A cheap model fed good signal beats an expensive model fed a list of names.
  • Following an explicit instruction turned out to separate the two models more than any factual error did. The prompt said not to force a savings framing on this account. Nova ignored that on four of five findings; Opus followed it, and additionally did cross-referencing that the flagging logic alone never asked either model to do.
  • Both scale-dependent bugs look identical to correct behavior until they don’t. Same root cause as the RSS bug: success and silent-partial-failure produce identical-looking output. And neither was found by watching something break — both came from deliberately asking “would this hold up bigger?” before handing the code to someone else’s AWS account.

How To Prevent It

Scenario What To Do
Writing an IAM policy for a Bedrock Cross-Region Inference Profile ID Don’t string-substitute the profile ID into a foundation-model ARN. Call GetInferenceProfile to see the real destination model ARNs and grant those explicitly (or wildcard the region on a specific model ID)
Regex-stripping XML namespaces to avoid namespace-aware parsing Strip the prefix off every tag occurrence too, not just the xmlns declarations on the root — otherwise the declarations disappear but the prefixed tags they were describing don’t
SAM template declares an explicit RoleName on an AWS::IAM::Role Use CAPABILITY_NAMED_IAM, not CAPABILITY_IAM
Any code path with a silent except: return [] for graceful degradation After a “successful” deploy, still read the actual logs of a real invoke — don’t just trust StatusCode: 200
Handing an LLM an account’s resource list and expecting useful advice Names alone produce advice that fits every account equally. Add usage-type or config-level granularity or the output degenerates to “consider configuring X” for anything on the list
Calling any AWS list/describe API expected to scale with account size Check the documented page-size cap and implement pagination even if the test account never reaches it — a truncated-but-valid response produces no error
Building a “top N for display” step on top of a detection step Run detection against the full, unfiltered list; truncate only at display time
Switching to a newer Claude generation on Bedrock Confirm it still accepts temperature, use an inference-profile model ID, and re-measure latency — adaptive thinking can multiply a single call’s runtime

Reference: