Infrastructure as Code and CI/CD on AWS
Infrastructure as code and continuous delivery solve the same operational problem from different angles: infrastructure should be reproducible, and changes should move through a controlled, testable release process. On AWS, CloudFormation and related tools define desired infrastructure, while pipeline services and deployment tools can validate, promote, and release those changes consistently.
The current DOP-C02 exam gives these skills major weight. SDLC Automation is 22 percent of scored content, while Configuration Management and Infrastructure as Code is another 17 percent. Candidates are expected to think beyond template syntax into testing, version control, deployment strategy, drift, rollback, monitoring, and secure automation.
A mature AWS delivery system should answer four questions: what state is intended, what change is proposed, how is it tested before production, and how can the team detect or reverse an unsafe deployment?
CloudFormation templates describe AWS resources and their properties as versioned code. AWS recommends storing templates in source control, using code review, linting, automated testing, and CI/CD rather than treating templates as files copied between laptops.
The AWS DevOps Engineer Professional certification is strongest when infrastructure changes follow the same engineering discipline as application code: reviewed, testable, traceable, and repeatable.
Manual console changes should be minimized because they create state that may not exist in the template. When emergency changes are required, reconcile them back into code quickly.
CloudFormation change sets preview which resources will be added, modified, replaced, or removed before the update is executed. This is especially important for stateful resources where a property change can cause replacement.
Review the proposed action, not only the template diff. A one-line change to a database name or immutable property can result in replacement and data loss if the architecture does not protect the resource.
The existing CloudFormation stack design is useful because large environments need structure as well as syntax.
Infrastructure drift occurs when a resource differs from the CloudFormation template because someone or something changed it outside the IaC workflow. AWS recommends regular drift detection and now supports drift-aware change sets for safer reconciliation.
Do not blindly overwrite drift. A change may have been made intentionally during an incident. Determine whether the template should be updated to reflect the new state or whether the resource should be returned to the declared configuration.
Drift is a governance signal: it shows where operational practice is bypassing the source of truth.
Continuous integration can run linting, unit tests, security checks, policy-as-code, template validation, and build steps every time infrastructure code changes. The goal is to catch defects while feedback is cheap.
Test in development or ephemeral environments before production. Validate that resources create successfully, permissions are correct, application dependencies can connect, and cleanup works.
The CloudFormation parameter-security pattern is a practical reminder that templates should not embed secrets simply because automation makes deployment convenient.
AWS CodePipeline can orchestrate source, build, test, approval, and deployment stages. The source can come from supported AWS or external providers, the build can use CodeBuild or another system, and deployment can target infrastructure or application services.
The CodePipeline workflow is most valuable when each stage has a clear gate. Source control proves what changed, build and tests validate it, approval protects high-risk environments, and deployment produces traceable output.
A pipeline should fail closed. If tests do not run or security validation is unavailable, the safest default is usually to stop rather than promote an unverified change.
In-place deployment can be simple but may expose users to a bad release. Rolling deployment reduces simultaneous impact but can create mixed-version behavior. Blue/green or canary patterns can reduce risk further when the application supports parallel environments or weighted traffic.
CodeDeploy and service-native deployment mechanisms can support these patterns depending on the workload. Define health checks, bake time, traffic shifting, and automatic rollback before production.
The adjacent SAA-C03 exam is useful architectural context because deployment strategy depends on load balancing, stateless design, data compatibility, and service resilience.
CI/CD systems often have permission to create or change large parts of the environment. Protect those identities through least privilege, scoped roles, approval boundaries, protected branches, and secure secret storage.
Do not embed access keys in build specifications or templates. Use IAM roles, Secrets Manager, Systems Manager Parameter Store, or supported dynamic references where appropriate.
A compromised pipeline can become a privileged infrastructure attack path, so source repository access and deployment roles deserve the same monitoring as production administrators.
CloudFormation can roll back failed stack updates, but application deployments and schema changes may require additional planning. Rolling back code is not enough if the new version already changed data in an incompatible way.
Use backward-compatible schema migrations, deployment order, feature flags, backups, and tested recovery procedures for stateful changes. Decide what happens when infrastructure succeeds but the application fails health checks afterward.
The DOP-C02 preparation material becomes more practical when every automation question includes failure and recovery.
For final practice, create one small CloudFormation stack in source control, add validation and tests, preview a change set, deploy through a pipeline, introduce drift, reconcile it, then perform a controlled failed deployment and rollback.
Capture CloudTrail activity and deployment logs so the team can answer who changed what and which pipeline run produced the current environment.
The goal of IaC and CI/CD is not maximum automation. It is a predictable change system where infrastructure state is reproducible, unsafe changes are stopped early, deployment evidence is preserved, and recovery does not depend on one engineer remembering what they clicked last week.
Stack organization affects maintainability. Very large templates can become difficult to review, while excessive fragmentation creates complicated dependencies between stacks. Use nested stacks, modules, StackSets, or separate service stacks where they create meaningful ownership boundaries. Outputs and parameters should expose only the interfaces other stacks need rather than coupling every resource together.
Multi-account delivery deserves deliberate design. Development, test, staging, and production are often separate AWS accounts, and the pipeline should assume roles into target accounts instead of storing long-lived credentials. Cross-account permissions, artifact encryption, KMS keys, and approval boundaries need to be planned before the pipeline is trusted with production deployment.
Policy as code can stop noncompliant infrastructure earlier in the lifecycle. CloudFormation Guard, organization policies, security scanners, custom tests, or CI checks can verify that resources meet encryption, logging, network, tagging, and access requirements before a change is promoted. Preventing a bad configuration at pull request time is cheaper than discovering it through a security incident.
Observability should cover the pipeline itself. Track failed builds, long queue times, repeated rollbacks, deployment duration, test failures, drift, and changes that bypass automation. A pipeline can be technically functional while users work around it because it is too slow or unreliable. Those workarounds are a signal to improve the delivery system.
Finally, practice immutable thinking where it fits. Replacing an application image or environment can be safer than modifying servers in place, especially for stateless workloads. The wider AWS certifications use this same operational principle across architecture, operations, security, and DevOps: known-good state should be reproducible instead of accumulating years of undocumented manual configuration.
StackSets become relevant when the same infrastructure or policy baseline must reach many accounts and Regions. Centralized deployment can improve consistency, but the change should still be tested against organizational units, permissions, and failure tolerance before broad rollout. Multi-account automation makes review and rollback even more important because one template can affect an entire organization.
The AWS CDK can define CloudFormation infrastructure using familiar programming languages, but the deployment still resolves into CloudFormation stacks and state. That means developers still need to understand change sets, replacement behavior, permissions, drift, and lifecycle. A higher-level language does not remove CloudFormation semantics.
Use promotion rather than rebuilding independently in each environment. The artifact or template tested in staging should be the one promoted to production, with environment-specific parameters supplied separately. Rebuilding from a moving branch at each stage can make “same release” mean different code.
Testing infrastructure code should include deletion and replacement behavior, not only creation. Tear down ephemeral environments, verify retention policies on stateful data, and confirm that cleanup does not leave orphaned resources or unexpectedly destroy protected data. Lifecycle quality includes safe removal as well as deployment.
Keep deployment evidence tied to the commit. A production stack should be traceable to the source revision, build, tests, approvals, and change set that created it. That audit trail shortens incident investigation and makes compliance reviews far easier.
Use tags and deployment metadata consistently so cost, ownership, and environment are visible after deployment. Infrastructure code should make those operational attributes repeatable rather than rely on manual cleanup later.
Keep a simple ownership map for each stack and pipeline so teams know who approves changes, who monitors failures, and who is responsible for reconciling drift after emergency work.