
What is the actual pricing mechanism behind serverless versus always-on compute?
Run Lambda at scale, say around 100 million invocations a month, and API Gateway plus duration charges can push the bill well past what you'd pay for a handful of t4g.medium instances doing the same job. Most serverless cost comparisons never get this far. They stop at the per-invocation rate and miss the crossover completely. So where does the math actually flip, and what should you be watching before it does?
Lambda pricing breaks into two pieces. There's a request charge, usually a small fraction of a cent per invocation, varying by region. And there's a duration charge based on memory allocated and milliseconds consumed, billed at rates like $0.0000166667 per GB-second. A function serving a 50ms response at 512MB costs a sliver of a cent per call. But that sliver multiplies linearly with volume. Unlike a container that can serve thousands of requests per second once it's up and running, there's no economy of scale here.
API Gateway adds a third layer that most cost comparisons skip entirely. Put Lambda functions behind API Gateway instead of a direct invocation path or an Application Load Balancer, and you're paying an extra $1.00 to $3.50 per million API calls, depending on whether you're running REST or HTTP APIs. Negligible at low volume. At high volume, it becomes the dominant line item. Any cost model that ignores API Gateway charges will understate the real bill once traffic scales, and that's exactly the gap in most comparisons: they price the function but not the path requests take to reach it.
Database infrastructure follows the same logic. Aurora Serverless v2 bills by Aurora Capacity Units at roughly $0.12 per ACU-hour, scaling automatically between a configured minimum and maximum. A provisioned Aurora instance like a db.r6g.large bills a fixed hourly rate no matter the load. Same crossover dynamic as Lambda versus EC2: serverless billing adapts to variable load, fixed billing rewards sustained, predictable load.
Cold starts complicate things further. A Lambda function that hasn't run recently takes extra time to spin up its execution environment before handling the first request. AWS Provisioned Concurrency fixes this by keeping instances warm, but it costs $0.0000041667 per GB-second whether or not requests show up. You're reintroducing an always-on cost into a pricing model built specifically to avoid always-on costs, which is worth sitting with before you decide Provisioned Concurrency is worth turning on.
Why does the crossover point land where it does, and what should you actually model?
The three cost layers above, request charges, duration charges, and API Gateway fees, stack on top of each other and determine where the crossover to containers actually happens. That point depends heavily on request volume and function duration. In my experience it tends to show up somewhere in the low millions of requests per month for typical workloads, though the exact number shifts hard based on how long each function runs. Short functions under 50ms stay cost-competitive with containers at much higher volumes than longer-running ones, because duration charges, not request charges, dominate the bill once invocation counts climb.
Past a certain volume, the math flips decisively toward containers. At very high invocation counts running behind API Gateway, the combined request, duration, and gateway charges can blow past what a small fleet of always-on instances behind an Application Load Balancer would cost for the same workload. The always-on setup can win by a wide margin, and it delivers more consistent latency too, since there's no cold start variance to worry about.
Provisioned Concurrency doesn't close this gap cheaply. Keeping multiple instances of a 512MB Lambda warm around the clock adds a recurring cost that scales with how many instances you keep warm, before any request or duration charges even get added on top. If your workload needs enough concurrency to justify a large number of pre-warmed instances, you're already approaching the volume where a small container fleet is cheaper and more stable on latency.
Aurora Serverless v2 plays by the same rules. A cluster sustaining a high average ACU load comparable to a db.r6g.large can cost noticeably more per month under Serverless v2 than the same instance running on-demand, and considerably more than under a Reserved Instance commitment. Serverless database pricing wins only when load genuinely swings, not when it sits at a consistent high plateau. A workload with meaningfully higher weekday traffic that drops off outside peak hours can average out to a modest savings under Serverless v2 compared to a provisioned instance sized for peak, and that savings is earned specifically because the load pattern is irregular rather than flat.
I/O charges under Aurora Standard add another variable. Per-request I/O pricing means a heavily loaded OLTP cluster can rack up a meaningful amount in I/O charges alone, on top of ACU costs, depending on volume. Check CloudWatch for VolumeReadIOPs and VolumeWriteIOPs before committing to Serverless v2. That tells you whether I/O-Optimized pricing would cut the bill further.
What does this mean for the architecture decision you're making right now?
So where crossovers happen for compute and databases is clear enough. The real question is which side of those crossovers your workload sits on, and what you do about it. If your traffic is bursty, irregular, or genuinely low-volume, serverless is still the right default, and the savings are real, not theoretical. A side project, an internal tool, an API that spikes unpredictably and sits idle most of the day: all of these avoid paying for 730 hours of capacity they never use. The lower end of the request-volume range isn't a hard ceiling, it's a planning threshold. Below it, model your actual function duration before assuming serverless wins. A 500ms function behaves very differently from a 20ms one at the same request volume.
For workloads already past the volume where containers get cheaper, or heading there within a planning horizon, switching to containers pays for itself quickly, not eventually. The gap between Lambda-behind-API-Gateway costs and equivalent container costs at high invocation volume isn't some edge case. It's the standard outcome once API Gateway charges stack on top of Lambda's per-invocation pricing. Run the numbers before you scale further. Running them after the bill arrives is too late to matter.
Database workloads deserve the same scrutiny Lambda gets. Aurora Serverless v2 can quietly cost noticeably more than a Reserved Instance once average utilization crosses a moderate-to-high share of equivalent provisioned capacity, because of how the ACU pricing model compounds under sustained load. If a workload looked bursty during prototyping but has since settled into consistent high load, that's a strong candidate for reverting to provisioned Aurora with a Reserved Instance commitment.
Separately: scheduling non-production environments to run only during working hours instead of 24/7 can meaningfully cut dev, test, and QA costs, and the size of the cut depends on how many hours you trim, regardless of which compute model you've chosen. You can do this before or alongside any serverless-versus-container decision, since it needs no re-architecting at all. Just schedule what's already running.
The question at the start was where the math flips and what to measure before it does. The answer is specific: track request volume, function duration, and whether API Gateway sits in the path. Those three numbers together determine which side of the crossover a workload lands on, not the per-invocation rate on its own. Model those before scaling, and the bill stops being a surprise.