Cloud Key Management Service Misconfiguration and Key Rotation Failures
Misconfigured KMS policies turn a single compromised credential into access to everything.

In January 2025, attackers running the Codefinger ransomware campaign compromised AWS credentials carrying nothing more than S3 read and write permissions. They encrypted victim data using SSE-C keys, a method where the key material never touches AWS's own storage, then demanded ransom. Because AWS never held a copy of those keys, there was no recovery path. Victims lost their data permanently because the attackers understood something most defenders had not fully priced in: when key material sits outside the platform's custody, there is no safety net underneath it.
That case points at something bigger than one ransomware group's clever use of SSE-C. Cloud key management services sit at the point where encryption keys for databases, object storage, secrets vaults, and inter-service traffic all come together. A mistake here does not behave like a mistake anywhere else in the stack. A misconfigured S3 bucket exposes one resource. A misconfigured KMS key policy exposes every resource that key protects, across every account and every workload allowed to use it.
The math is unforgiving here. A key policy that grants Principal: * with no conditions attached is functionally public. Any AWS principal, including ones sitting in entirely separate accounts, can use it. Think of a public S3 bucket, except this one holds the keys that unlock everything else built to protect that bucket. Treating KMS misconfiguration as routine cloud hygiene misses the point. It is a structural failure in the one control that every other data protection measure depends on.
The three structural patterns that cause nearly every KMS failure
Looking across documented KMS incidents, the failures sort into three repeatable patterns: access that is too broad, rotation that is handled by hand instead of by automation, and key material that has leaked out of the KMS and into code or pipeline configuration.
A hardcoded key sitting in a CI/CD pipeline, one that has never been rotated and carries wildcard permissions, is all three failures stacked on top of each other at once. When that happens, the blast radius combines the damage of all three patterns together.
Key management rarely collapses in one dramatic event. It decays slowly, through ownership that nobody quite owns, telemetry that nobody's watching closely enough, and workarounds that start as temporary fixes and end up as how things are just done. Naming these three patterns turns that slow decay into something a team can actually see and act on, rather than a vague sense that things feel less buttoned-up than they used to.
The next three sections take each pattern in turn: how it works, a real incident that shows it in action, and what closes the gap. Read together, they form a diagnostic framework, a way to ask where a given environment actually stands.
Pattern one: overprivileged access that turns one compromised identity into a full-environment breach
Overprivileged KMS access takes any compromised identity, a stolen session token, a misconfigured service account, an insider acting in bad faith, and turns it into a master key for everything that KMS key touches. The identity itself barely matters once it's compromised. What matters is the ceiling on what that identity is allowed to do.
Wildcard action or resource policies rank among the most dangerous IAM misconfigurations in cloud environments today. A single compromised credential carrying broad permissions can spin up new admin accounts, pull data out of every S3 bucket in reach, or deploy workloads of the attacker's choosing. None of that requires any additional cleverness on the attacker's part. The permission to do it was already sitting there, granted in advance.
The TeamPCP campaign, which also touched LiteLLM and Aqua Security in March 2026, shows this escalation path in documented detail. Attackers stole CI/CD credentials by poisoning security-scanning tools that victim pipelines trusted, harvested AWS and other cloud credentials from those compromised pipelines, and used the stolen secrets to move laterally across Kubernetes clusters and cloud environments. None of that lateral movement would have gone anywhere if the credentials in question had been scoped tightly to what each pipeline actually needed.
Google Cloud's H1 2026 report describes a related pattern worth sitting next to the TeamPCP case. Attackers abused a trust relationship connecting GitHub to AWS through OpenID Connect, then used an overly permissive cloud role to create administrative access for themselves. The entry point in that intrusion was a trust boundary that had been configured too loosely, not a brute-force attack against a password or a key.
Cross-account trust deserves its own attention inside the KMS conversation specifically. A role trust policy that allows any principal from an external account, with no condition constraints attached, creates a door nobody's watching. The "confused deputy" problem compounds this: an attacker tricks a trusted, more privileged service into misusing its own authority on the attacker's behalf, so the privilege path is never visible as a direct grant to the attacker. It hides inside a service that was never supposed to be abusable.
Audit frameworks require separating key administration from workload administration for good reason, and teams routinely blur that line anyway. When that separation disappears, overbroad access becomes one of the clearest signs that key management controls have stopped functioning, and auditors go looking for it.
Closing this gap starts with scoping KMS key policies down to named principals and specific actions, and removing any Principal: * grant that doesn't carry a compensating condition alongside it. AWS IAM Access Analyzer can generate findings for external access to KMS keys, though it won't surface privilege escalation paths on its own; tools like Cloudsplaining or PMapper fill that gap by mapping out how a given identity could escalate its own access. Every compute resource carrying an instance role that touches sensitive keys should run IMDS v2, no exceptions. And cross-account trust policies need a regular audit for missing condition constraints, paired with CloudTrail logging and alerting on sts:AssumeRole events coming from unusual source accounts or unfamiliar IPs.
Pattern two: manual rotation processes that fail under the conditions that make rotation most necessary
Manual rotation doesn't fail because engineers get careless. It fails because the conditions that make rotation urgent in the first place, incident response, a late-night on-call page, pressure to get production back online, are the exact same conditions under which a multi-step manual process is most likely to skip a verification step.
Cloudflare's R2 outage on March 21, 2025 is the clearest documented case of this. A credential rotation was kicked off as planned. The new credential got deployed, but it landed in the dev environment instead of production, because the --env production parameter was left off the command that should have carried it. The team then deleted the old credentials, working from the assumption that production had already picked up the new ones, but it hadn't. Production kept trying to authenticate with credentials that no longer existed, and the result was total write failures and partial read failures spreading globally for over an hour.
Cloudflare's own incident report named two structural gaps that caused that sequence: no real-time visibility into which credentials were actually in active use at the moment of deletion, and no verification step built in before the old keys got deleted. The rotation process existed on paper. It just wasn't complete in practice, and the gap between those two things is where the outage lived.
Microsoft's Exchange key incident in 2023 shows the opposite failure mode, sitting at the opposite end of the same problem as Cloudflare's case. Manual rotation was stopped after a 2021 cloud outage that the rotation process itself had caused. Rather than fixing the process or automating it, Microsoft never resumed rotation and never built an automated replacement, a decision that eventually fed into a large-scale compromise. Fear of the process that was supposed to protect the system became, on its own, a security liability.
The TeamPCP compromise adds a third angle to this same pattern: rotation that happens but doesn't finish. The attacker group retained access after a credential rotation that was incomplete. The claim "we rotated the keys" needs to be backed by evidence that the rotation actually completed end to end, not just a status message confirming the job ran.
One gap in this pattern gets missed even by teams doing everything right at the KMS layer itself. By default, BigQuery does not automatically rotate a table's encryption key when the associated Cloud KMS key rotates. Existing tables keep using whatever key version they were created with, indefinitely, while the KMS console shows the rotation as complete. A compliance-focused team can audit the KMS key itself, confirm the rotation happened on schedule, and still have tables sitting on encryption keys that are years old, because the console's confirmation and the downstream service's actual behavior are two different things.
Fixing this starts with automating the verification of deployed credentials before anyone deletes the old ones. Automated checks that confirm consumption context catch environment mismatches exactly like the one Cloudflare hit. Rolling rotations help too: spin up the new credential, confirm it's actually being used, then decommission the old one, so a failure in the new credential still leaves production safe and reversible. Teams need a complete map of which workloads consume which key versions before any rotation job runs, and every rotation needs a follow-up audit of downstream service behavior. A green checkmark in the KMS console confirms the key rotated. It says nothing about whether every service depending on that key actually re-encrypted with it.
Pattern three: keys embedded in code and the limits of detection-only responses
Keys hardcoded into source code or CI/CD configuration are the predictable output of any workflow that asks developers to pass credentials between systems by hand, with no secrets manager sitting in the middle to mediate that handoff. And once a key is found sitting in a repository, detecting it is only half the job. Without revoking it, the exposure stays live.
The CI/CD pipeline usually carries more privilege than any other system in a given infrastructure, which makes it the highest-value target for this exact failure. Hardcoded secrets in pipeline YAML or application source are the most common way credentials end up exposed, and removing a secret from the current commit does nothing to remove it from git history. Anyone who can read the repository's full history can still find it. Once a key has been exposed this way, the only safe assumption is that it's been compromised, and the response is rotation, immediately, not a quiet cleanup commit.
Secrets sprawl compounds the exposure as time passes. A large share of secrets leaked years ago, 64% of valid secrets from 2022, are still active and exploitable today. That number says something specific: scanning for exposed secrets without a governance process to revoke them leaves the exposure sitting there indefinitely, found but still usable.
Stale credentials carrying admin-level permissions are a prime target for credential stuffing and leaked-secret exploitation, precisely because nobody's watching whether they're still in use. AWS IAM Access Analyzer's Unused Access Analyzer addresses this directly, showing which permissions haven't been used within a configurable tracking period, which surfaces dormant keys that should have been revoked long ago.
The response to finding a hardcoded key in a codebase can't stop at deleting the line it appears on. The key needs to be treated as compromised and rotated right away, and the rest of that repository's history needs an audit for any other exposures sitting alongside it. Replacing hardcoded credentials with references to a secrets manager, AWS Secrets Manager, HashiCorp Vault, or GCP Secret Manager among them, keeps credentials from ever being written into source or pipeline configuration. Secret-scanning tools, whether that's native GitHub secret scanning or AWS IAM Access Analyzer's policy validation running inside CI/CD, need to run against every commit and against the full historical record, not just whatever sits at HEAD. And wherever possible, CI/CD pipelines should authenticate using short-lived credentials through OIDC-based identity federation, so there's no long-lived secret sitting in pipeline configuration for an attacker to find.


