Skip to main content

Posts

Least-Privilege IAM and Secret Handling as Architecture Decisions, Not Compliance Cleanup

When I review cloud architecture in modernization projects, I often find IAM policies and secret-handling patterns treated as post-deployment security tasks rather than foundational design decisions. Teams grant broad permissions during development, defer secret rotation until a compliance audit surfaces the risk, and treat least-privilege access as a hardening step that happens after the application works. This approach consistently produces architectures that are expensive to secure later and difficult to operate safely at scale. I learned this lesson the hard way. In one AWS data pipeline, early prototypes used a single service role with wide S3, Redshift, and Lambda permissions. The role worked across environments, simplified local testing, and allowed rapid iteration. When the pipeline moved to production and began processing regulated data, tightening those permissions required rewriting deployment scripts, breaking existing integrations, and coordinating changes across multip...
Recent posts

Building Resilient Data Connectors When External APIs Fail, Drift, or Throttle

External APIs break in predictable ways. They rate-limit your requests during peak hours. They change field names without warning. They return HTTP 200 with an error object nested three levels deep in JSON. They timeout after exactly 29 seconds, one second before your own timeout fires. If your data connector treats these as exceptional cases, your pipeline stops working the day your business depends on it. I have built data connectors for regulated environments where the external API was a legacy SOAP service, a partner REST endpoint with unpredictable availability, and a SaaS vendor that changed response schemas without versioning. The lesson from all three is the same: resilience is not a retry loop wrapped around an HTTP client. It is a design decision about failure ownership, state management, and operational visibility. Failure Modes You Must Design For, Not Catch Later Most data connector failures fall into three categories: transient errors you can retry, permanent error...

Why CI/CD Still Fails When the Pipeline Turns Green

A green build does not mean your deployment worked. I learned this the hard way while modernizing a legacy .NET system that had spent years accumulating manual release steps outside the automation. The CI/CD pipeline would pass, the artifact would land in the environment, and then someone would remember the configuration flag that never made it into the repository, or the database migration script that lived in a wiki, or the API key rotation that happened through a support ticket. The pipeline said success, but the application was broken in production. The problem was not the CI/CD tool. We were using GitHub Actions, which is capable and flexible. The problem was that CI/CD had been treated as a checkbox—something you set up once to automate the build and maybe run a few tests—rather than a system designed to make deployments smaller, safer, and repeatable. The hidden manual steps were not documented anywhere the pipeline could see them, so they became failure points every time we ...

AWS Lambda ARM64 Container Images: The Runtime Decision I Wish I Had Made Earlier

The blog automation agent I run in AWS Lambda uses an ARM64 container image. The decision to switch from the managed Python runtime to a containerized deployment came after the service had been running reliably for months, and the migration itself was prompted by a single requirement: runtime credential management became complex enough that baking dependencies into a container image was simpler than coordinating layer versions and environment configuration. What I did not expect was how much the choice of ARM64 versus x86_64 would matter once the container deployment model was in place. This is not an article about containers versus managed runtimes. That choice depends on your deployment pipeline, dependency management, and operational boundaries, and I covered the broader execution shape decision in an earlier post. This is about the narrower question: once you have decided to deploy a Lambda function as a container image, should you build for ARM64 or x86_64? The answer is less o...

The Production Cost Traps Behind VPC Networking, NAT Gateways, and Data Transfer

I spent three months optimizing an AWS data pipeline that processed billions of records through Apache Spark and loaded them into Amazon Redshift. The job ran reliably. Query performance was acceptable. The engineering team had moved on to the next feature. Then the monthly bill arrived, and a single line item stopped the conversation: $4,200 for NAT Gateway data processing charges. The pipeline itself was efficient. We had tuned the Spark partitions, compressed the intermediate files, and designed the Redshift distribution keys to minimize broadcast joins. But we had placed every Lambda function, every Spark worker, and every API client inside a VPC with private subnets, and we had routed all outbound traffic through a NAT Gateway in each availability zone. The architecture diagram looked secure. The cost structure was unsustainable. VPC networking decisions feel like infrastructure hygiene when you make them. They become budget traps when you forget that AWS charges for data mov...

Designing Tenant Isolation Before the Data Model Becomes Expensive to Change

You're building a SaaS product or internal platform that serves multiple customers, business units, or teams. Right now, you have a single database, a shared schema, and a WHERE clause filtering on tenant_id. It works. The question is: will it still work when you need to move a high-value customer to dedicated infrastructure, when compliance requires physical separation, or when one tenant's query load starts degrading performance for everyone else? Tenant isolation is not just a security concern. It's an architectural decision that shapes your data model, your query patterns, your deployment topology, and your operational flexibility. Make the wrong choice early, and you'll face expensive migrations under pressure. Make the right choice, and you preserve options without over-engineering. Isolation Models and Their Tradeoffs There are three common isolation models: shared schema with row-level filtering, schema-per-tenant, and database-per-tenant. Each has a diff...

When to Split a Microservice and When to Keep It Together

You have a monolithic .NET service handling customer orders, inventory checks, and email notifications. A colleague suggests splitting it into three microservices. Another argues the split will triple operational overhead without delivering business value. Both are probably right, but neither answer helps you decide what to do Monday morning. After modernizing multiple legacy .NET systems and operating microservices architectures on both AWS and Azure, I've learned that the decision to split a service is less about embracing distributed systems and more about matching architectural boundaries to the operational consequences your team can actually handle. The wrong split creates alert fatigue, deployment dependencies, and debugging sessions that span five repositories. The right boundary reduces coupling, enables independent deployments, and makes production incidents easier to isolate. Start With the Deployment Boundary, Not the Domain Model The first question is not whether...