AWS Lambda vs. Google Cloud Functions: Running AI Inference Workloads
Last month, two clients came to me with nearly identical requests. They wanted to run lightweight AI models in a serverless setup. One was already deeply tied into the AWS ecosystem, and the other was running most of its pipeline on GCP. That gave me the rare opportunity to work with both platforms almost simultaneously. I can still picture that Saturday morning in a coffee shop in Decatur, two laptops open, switching back and forth between both consoles.
The comparison was between AWS Lambda and Google Cloud Functions. The scenario for both was deploying a small text classification model (ONNX format, roughly 45MB) and exposing it as a REST endpoint.
Cold Start
When you talk about AI inference in a serverless context, cold start is always the first thing that comes up. Think of it like starting a car that's been sitting in the garage all winter. You have to wait for the engine to warm up before you can drive.
On the Lambda side, I bundled the model file as a layer with the Python runtime. Cold starts frequently ran over 5 seconds, but bumping the memory from 1,024MB to 3,008MB noticeably reduced that. Turning on Provisioned Concurrency can virtually eliminate cold starts, but at that point the word "serverless" starts to lose its meaning. Costs go up, too.
On the Cloud Functions side (2nd gen, based on Cloud Run), setting the minimum instance count to 1 achieved a similar effect. The cold start itself felt slightly longer than Lambda's, but after warming up, response times on consecutive calls were more stable. This isn't a formal benchmark. It's what I observed while going back and forth between the two projects.

Pricing Structure
Both charge based on the number of invocations and execution time (in GB-second units), so the basic structure is the same. The differences are in the details.
Lambda offers billing in 1ms increments. This works in your favor for short inference calls. On the other hand, if you use Provisioned Concurrency, you incur costs even during idle time, so you need to crunch the numbers carefully for workloads with irregular traffic.
Cloud Functions has also moved from 100ms increments to more granular billing, but the cost of maintaining minimum instances added up surprisingly fast. One client had so little monthly traffic that the minimum instance maintenance cost exceeded the actual invocation cost. You need to map out your traffic patterns first and then decide which option makes sense.
Developer Experience
This is where the two platforms really diverge in character.
Lambda has a broad ecosystem. There are plenty of deployment tools (SAM, CDK, Serverless Framework, and more) along with abundant community examples. However, when deploying AI models, you have to deal with layer size limits (250MB uncompressed) and packaging issues yourself. Things have gotten much better since container image support was introduced, but the initial setup still requires some time wrestling with configuration.
Cloud Functions 2nd gen essentially sits on top of Cloud Run, so container-based deployment feels natural. For workloads with heavy dependencies like AI models, this structure was convenient. You just build a Docker image and push it. On the other hand, the local emulator experience was more mature on the Lambda side.

Which Did We Choose
The AWS client went with Lambda in container image mode. Their VPC, IAM, and S3 pipelines were all already built around Lambda, so there was no reason to switch platforms. The GCP client went with Cloud Functions 2nd gen, but we set up the architecture to share container images so they could move to Cloud Run without friction as traffic grows.
There was no single answer to "which platform is better." For both projects, the existing infrastructure context and the tools the team was comfortable with determined the choice. If you pick based solely on a technical spec comparison chart, you'll miss a lot. Which console the team opens every day, and where their existing CI/CD is hooked in, those turned out to be the more important variables.
Comments