CodamAIDocs
Topicdone

When CIAS or Keycloak fails

What happens during an outage: the tenant gate answers from memory, registration and administration answer with 503, a failed provisioning stays repeatable.

Variants
CIAS downKeycloak down during a requestKeycloak down during registrationKeycloak down during administrationProvisioning failedRestart during an outage

What this is about

CIAS and Keycloak are two different services, and either can fail at some point. What happens then depends on who is asking whom at that moment:

  • CIAS answers, on every request to CDMS, the question “is this tenant served?” and, if set up, “which attribute values does the person have in this tenant?”.
  • Keycloak issues tokens, exchanges them and takes every change to accounts, roles and organizations.

Asking both questions live on every request would be slow and fragile. So the filter chain remembers the answers for a short time. During an outage, exactly this memory decides who can keep working.

Who asks whom and when

flowchart LR
    C["Client"] --> F["Filter chain"]
    F -- "keys for the signature check<br/>(remembered for a while)" --> K["Keycloak"]
    F -- "exchange the token<br/>(up to 5 min per token)" --> K
    F -- "tenant served?<br/>(remembered 30 s)" --> CI["CIAS"]
    F -- "attribute values in the tenant<br/>(remembered 30 s)" --> CI
    R["Registration,<br/>administration"] -- "live, nothing remembered" --> K

Read it like this: the filter chain needs Keycloak and CIAS on every request, but remembers the answers. Registration and administration, on the other hand, talk to Keycloak live, because they change something.

Stepasksremembersduring an outage
Check the token signatureKeycloak, for the public keys of the realmthe keys, for a whileAs long as the keys are there, the check does not need Keycloak. If it needs new ones and gets none, the request fails.
Exchange the tokenKeycloakthe exchanged token, up to 5 minutes per token (codamai.cias.token-exchange.ttl), never longer than it is validA token that was already exchanged keeps working. For a new token see below.
Admit the tenantCIASevery answer for 30 seconds (codamai.cias.tenant-gate.ttl)the last answer, as long as it is at most 15 minutes old (codamai.cias.tenant-gate.stale-ceiling), otherwise 403
Attribute values per tenantCIASevery answer for 30 seconds (codamai.cias.attribute-lookup.ttl)the last values, as long as they are at most 15 minutes old (codamai.cias.attribute-lookup.stale-ceiling), otherwise 403
Log in, renew the tokenKeycloaknothingnot possible

CIAS down

If CIAS fails, the filter chain works with what it has remembered. The rule is the same for the tenant gate and for the attribute values:

Tenant gate and attribute values when CIAS does not answer
Remembered answerResult
present, at most 15 minutes oldthe remembered answer still applies, even if it was no
present, but older403 cias.authentication.tenant-not-served
none403 cias.authentication.tenant-not-served

In one sentence: a CIAS outage does not throw out anyone who was already working, and does not let anyone new in. A tenant that was last suspended stays suspended. A person keeps the values they last worked with, and gets no new ones.

But not indefinitely. The 15 minutes are the staleness limit: they bridge a restart or a rollout without refusing anybody. If the outage lasts longer, the filter chain stops answering from memory and refuses. The reason is simple: while it answers from memory it learns nothing new — a suspension issued during that time would otherwise never arrive.

What “does not answer” means depends on the operating mode:

embeddedstandalone
How the question is askedmethod call in the same processHTTP request with a service token
What counts as an outagethe lookup throws an error, for example because the system database cannot be reachedno connection after 2 seconds, no answer after 2 seconds, an error code, an unreadable answer
What does not count as an outage–404: CIAS does not know the tenant, the gate refuses and remembers that
What refuses immediately–a rejected service token (401 or 403): the request is refused even if a matching answer is remembered

The service token is a special case, because it does not heal by itself. Running standalone, CDMS identifies itself to CIAS with a fixed token. If CIAS rejects it, the token has expired or is set wrongly, and every further question fails the same way — until somebody replaces it. That is why this refuses instead of bridging: an outage that does not pass would otherwise be answered from memory forever.

The details of the gate are in Admit the tenant (tenant gate). The same applies to work without a request, for example a timer running for a tenant, see Working for a tenant without a request.

Keycloak down during a request

A request to CDMS while Keycloak does not answer

When: The person is working right now, their token was already exchanged in the last few minutes.

The filter chain takes the exchanged token from its memory and does not ask Keycloak. The request runs completely normally as long as the token and the remembered exchange are valid.

Result: The request reaches the application.

When: A new token has to be exchanged, Keycloak answers, but with an error code.

CIAS writes an error to the log and lets the request continue without an identity. The application sees no person and no roles.

Result: Every role check refuses, usually with 403. See Token exchange.

When: A new token has to be exchanged, but Keycloak cannot be reached at all.

The exchange fails before there is an answer.

Result: The request fails.

When: The person logs in again, or their token expires and has to be renewed.

Only Keycloak issues tokens. Without Keycloak there is no new token.

Result: The person can only continue once Keycloak is back. See Renew the token.

The tenant gate does not depend on Keycloak. It asks CIAS, not Keycloak.

Keycloak down during registration

Registration talks to Keycloak live in several places. If Keycloak cannot be reached, CIAS answers with 503 cias.iam.unavailable. What is left afterwards depends on the step:

Where registration meets Keycloak

When: POST /cias/registration/self, or an administrator creates or invites a person.

CIAS first asks Keycloak whether the address already exists. If that fails, nothing has been saved yet and no mail has been sent.

Result: 503. The person submits the form again later.

When: The person clicks the confirmation link, and provisioning starts right away.

Provisioning fails because of Keycloak. CIAS sets the registration to FAILED and saves that. A second click on the same link answers with 202, but does not continue the provisioning: the link has already been redeemed.

Result: 503 on the first click. It only continues with retry by a platform administrator.

When: An administrator approves a registration, activates it without a link, or calls retry.

Here too, provisioning starts and fails because of Keycloak. Afterwards the registration is FAILED.

Result: 503. Call retry later.

Every save of the registration is a small step of its own. That is why the state FAILED stays saved even when the request ends with an error. A platform administrator finds the registration in the list with state=FAILED. See What happens on completion.

Keycloak down during administration

Whoever changes permissions writes to CIAS and to Keycloak. If Keycloak fails in between, half a step is left over. CIAS chooses the order so that this half step always leaves fewer permissions, never more: when taking away, it writes to Keycloak first; when granting, it writes to CIAS first.

Administration while Keycloak cannot be reached
CallDirectionAnswer and what is left over
grant a rolegrants503. The grant is in CIAS, it is still missing from the token. Repeat the call.
revoke a roletakes away503. Nothing changed, the role still applies. Repeat the call.
suspend or close an accounttakes away503. Nothing changed. Repeat the call.
reactivate an accountgrants503. The record is active, the account in Keycloak is still disabled. Repeat the call.
create a group, add roles or membersgrantsSuccess. The group is PENDING in CIAS, the reconciliation carries it to Keycloak later.
take roles or members out of a group, delete a grouptakes away503. Nothing changed. Repeat the call.
a time-limited role expires (timer)takes awayThe run skips the grant and carries on with the others. The next run tries again.

Repeating is safe in every case. A second grant of the same role does not create a second record, but writes the existing one to Keycloak again. Why the order is exactly this way: The write order.

Provisioning failed

Some processes need several steps, and an outage can hit them in the middle. CIAS then leaves them in a state from which you can start again. None of these states lets anyone work who should not.

What failsState afterwardsIs anyone working already?How it continues
Provisioning of a registrationregistration FAILEDnoPOST /cias/admin/registrations/{id}/retry, or discard if it no longer makes sense
Setting up a new tenanttenant with rollout FAILED, standing PENDING, answer 502no, the gate does not admit itPOST /cias/admin/tenants/{id}/retry-provisioning
Setup stops without reporting backrollout stays IN_PROGRESSnothe reconciliation sets it to FAILED after the deadline, then retry. See Create and provision a tenant
Group does not reach Keycloakgroup PENDINGthe permissions from the group do not apply yetgroup reconciliation, see Reconciliation with Keycloak
Role reconciliation at startup does not reach Keycloaknothing written, message in the log–The application starts anyway. Reconcile again later.

Restart during an outage

The memory of the filter chain lives only in memory. When an application restarts, it is empty:

  • If CIAS is still down, the tenant gate does not know a single tenant and refuses all requests with a tenant until CIAS answers again.
  • If Keycloak is still down, every token has to be exchanged anew, and that does not work. The keys for the signature check are missing too.

So a restart “to be safe” makes an outage worse, not better.

Pitfalls

Next

Sources in the code and the knowledge base
  • CIAS/cias-authentication – TenantGate.admit (cache, outage rule), AttributeLookup (same outage rule), TokenExchangeService (jti cache, call, refusal), TokenExchangeProperties (ttl 5 min), TokenParser.tokenParser, JwtDecoderUtil
  • CIAS/cias-tenancy-client – RemoteTenantLookupAdapter.find (404, 401/403, other codes), CiasTenancyClientProperties (timeouts 2 s)
  • CIAS/cias-iam-keycloak – KeycloakAdminApi (ResourceAccessException and 5xx → IamUnavailableException)
  • CIAS/cias-registration – RegistrationService (register, verify, provision, retry), Registration (verify, startProvisioning, fail), RegistrationController (constant body), RegistrationExceptionHandler.iamDown, JpaRegistrationRepositoryAdapter (every save its own transaction)
  • CIAS/cias-authorization – RoleAssignmentService (grant: record first; revoke: Keycloak first; synchronizeDueAssignments), GroupService (project, tolerate, restrict), RoleReconciliationService.execute, AuthorizationExceptionHandler.providerUnavailable
  • CIAS/cias-user – UserService (suspend, close, activate), UserExceptionHandler.providerUnavailable
  • CIAS/cias-tenancy – TenantService (rollOut, retryProvisioning), Tenant.isServedOn, TenantExceptionHandler.provisioning
  • CIAS/cias-spring-boot-starter – CiasAutoConfiguration.ciasRoleReconciliationRunner, TenantReconciliationScheduler
  • CIAS/cias-kernel/docs/adr – ADR-022 §4, §5; CIAS/cias-authorization/docs/adr – ADR-034
Search