What this is about
CIAS and Keycloak are two different services, and either can fail at some point. What happens then depends on who is asking whom at that moment:
- CIAS answers, on every request to CDMS, the question “is this tenant served?” and, if set up, “which attribute values does the person have in this tenant?”.
- Keycloak issues tokens, exchanges them and takes every change to accounts, roles and organizations.
Asking both questions live on every request would be slow and fragile. So the filter chain remembers the answers for a short time. During an outage, exactly this memory decides who can keep working.
Who asks whom and when
flowchart LR
C["Client"] --> F["Filter chain"]
F -- "keys for the signature check<br/>(remembered for a while)" --> K["Keycloak"]
F -- "exchange the token<br/>(up to 5 min per token)" --> K
F -- "tenant served?<br/>(remembered 30 s)" --> CI["CIAS"]
F -- "attribute values in the tenant<br/>(remembered 30 s)" --> CI
R["Registration,<br/>administration"] -- "live, nothing remembered" --> K
Read it like this: the filter chain needs Keycloak and CIAS on every request, but remembers the answers. Registration and administration, on the other hand, talk to Keycloak live, because they change something.
| Step | asks | remembers | during an outage |
|---|---|---|---|
| Check the token signature | Keycloak, for the public keys of the realm | the keys, for a while | As long as the keys are there, the check does not need Keycloak. If it needs new ones and gets none, the request fails. |
| Exchange the token | Keycloak | the exchanged token, up to 5 minutes per token (codamai.cias.token-exchange.ttl), never longer than it is valid | A token that was already exchanged keeps working. For a new token see below. |
| Admit the tenant | CIAS | every answer for 30 seconds (codamai.cias.tenant-gate.ttl) | the last answer, as long as it is at most 15 minutes old (codamai.cias.tenant-gate.stale-ceiling), otherwise 403 |
| Attribute values per tenant | CIAS | every answer for 30 seconds (codamai.cias.attribute-lookup.ttl) | the last values, as long as they are at most 15 minutes old (codamai.cias.attribute-lookup.stale-ceiling), otherwise 403 |
| Log in, renew the token | Keycloak | nothing | not possible |
CIAS down
If CIAS fails, the filter chain works with what it has remembered. The rule is the same for the tenant gate and for the attribute values:
| Remembered answer | Result |
|---|---|
| present, at most 15 minutes old | the remembered answer still applies, even if it was no |
| present, but older | 403 cias.authentication.tenant-not-served |
| none | 403 cias.authentication.tenant-not-served |
In one sentence: a CIAS outage does not throw out anyone who was already working, and does not let anyone new in. A tenant that was last suspended stays suspended. A person keeps the values they last worked with, and gets no new ones.
But not indefinitely. The 15 minutes are the staleness limit: they bridge a restart or a rollout without refusing anybody. If the outage lasts longer, the filter chain stops answering from memory and refuses. The reason is simple: while it answers from memory it learns nothing new — a suspension issued during that time would otherwise never arrive.
What “does not answer” means depends on the operating mode:
| embedded | standalone | |
|---|---|---|
| How the question is asked | method call in the same process | HTTP request with a service token |
| What counts as an outage | the lookup throws an error, for example because the system database cannot be reached | no connection after 2 seconds, no answer after 2 seconds, an error code, an unreadable answer |
| What does not count as an outage | – | 404: CIAS does not know the tenant, the gate refuses and remembers that |
| What refuses immediately | – | a rejected service token (401 or 403): the request is refused even if a matching answer is remembered |
The service token is a special case, because it does not heal by itself. Running standalone, CDMS identifies itself to CIAS with a fixed token. If CIAS rejects it, the token has expired or is set wrongly, and every further question fails the same way — until somebody replaces it. That is why this refuses instead of bridging: an outage that does not pass would otherwise be answered from memory forever.
The details of the gate are in Admit the tenant (tenant gate). The same applies to work without a request, for example a timer running for a tenant, see Working for a tenant without a request.
Keycloak down during a request
When: The person is working right now, their token was already exchanged in the last few minutes.
The filter chain takes the exchanged token from its memory and does not ask Keycloak. The request runs completely normally as long as the token and the remembered exchange are valid.
Result: The request reaches the application.
When: A new token has to be exchanged, Keycloak answers, but with an error code.
CIAS writes an error to the log and lets the request continue without an identity. The application sees no person and no roles.
Result: Every role check refuses, usually with 403. See Token exchange.
When: A new token has to be exchanged, but Keycloak cannot be reached at all.
The exchange fails before there is an answer.
Result: The request fails.
When: The person logs in again, or their token expires and has to be renewed.
Only Keycloak issues tokens. Without Keycloak there is no new token.
Result: The person can only continue once Keycloak is back. See Renew the token.
The tenant gate does not depend on Keycloak. It asks CIAS, not Keycloak.
Keycloak down during registration
Registration talks to Keycloak live in several places. If Keycloak cannot be reached, CIAS answers with 503 cias.iam.unavailable. What is left afterwards depends on the step:
When: POST /cias/registration/self, or an administrator creates or invites a person.
CIAS first asks Keycloak whether the address already exists. If that fails, nothing has been saved yet and no mail has been sent.
Result: 503. The person submits the form again later.
When: The person clicks the confirmation link, and provisioning starts right away.
Provisioning fails because of Keycloak. CIAS sets the registration to FAILED and saves that. A second click on the same link answers with 202, but does not continue the provisioning: the link has already been redeemed.
Result: 503 on the first click. It only continues with retry by a platform administrator.
When: An administrator approves a registration, activates it without a link, or calls retry.
Here too, provisioning starts and fails because of Keycloak. Afterwards the registration is FAILED.
Result: 503. Call retry later.
Every save of the registration is a small step of its own. That is why the state FAILED stays saved even when the request ends with an error. A platform administrator finds the registration in the list with state=FAILED. See What happens on completion.
Keycloak down during administration
Whoever changes permissions writes to CIAS and to Keycloak. If Keycloak fails in between, half a step is left over. CIAS chooses the order so that this half step always leaves fewer permissions, never more: when taking away, it writes to Keycloak first; when granting, it writes to CIAS first.
| Call | Direction | Answer and what is left over |
|---|---|---|
| grant a role | grants | 503. The grant is in CIAS, it is still missing from the token. Repeat the call. |
| revoke a role | takes away | 503. Nothing changed, the role still applies. Repeat the call. |
| suspend or close an account | takes away | 503. Nothing changed. Repeat the call. |
| reactivate an account | grants | 503. The record is active, the account in Keycloak is still disabled. Repeat the call. |
| create a group, add roles or members | grants | Success. The group is PENDING in CIAS, the reconciliation carries it to Keycloak later. |
| take roles or members out of a group, delete a group | takes away | 503. Nothing changed. Repeat the call. |
| a time-limited role expires (timer) | takes away | The run skips the grant and carries on with the others. The next run tries again. |
Repeating is safe in every case. A second grant of the same role does not create a second record, but writes the existing one to Keycloak again. Why the order is exactly this way: The write order.
Provisioning failed
Some processes need several steps, and an outage can hit them in the middle. CIAS then leaves them in a state from which you can start again. None of these states lets anyone work who should not.
| What fails | State afterwards | Is anyone working already? | How it continues |
|---|---|---|---|
| Provisioning of a registration | registration FAILED | no | POST /cias/admin/registrations/{id}/retry, or discard if it no longer makes sense |
| Setting up a new tenant | tenant with rollout FAILED, standing PENDING, answer 502 | no, the gate does not admit it | POST /cias/admin/tenants/{id}/retry-provisioning |
| Setup stops without reporting back | rollout stays IN_PROGRESS | no | the reconciliation sets it to FAILED after the deadline, then retry. See Create and provision a tenant |
| Group does not reach Keycloak | group PENDING | the permissions from the group do not apply yet | group reconciliation, see Reconciliation with Keycloak |
| Role reconciliation at startup does not reach Keycloak | nothing written, message in the log | – | The application starts anyway. Reconcile again later. |
Restart during an outage
The memory of the filter chain lives only in memory. When an application restarts, it is empty:
- If CIAS is still down, the tenant gate does not know a single tenant and refuses all requests with a tenant until CIAS answers again.
- If Keycloak is still down, every token has to be exchanged anew, and that does not work. The keys for the signature check are missing too.
So a restart “to be safe” makes an outage worse, not better.