ECPDS Plugin Runbook
Use this page to investigate ECPDS authorization problems on watch or replay requests. For setup, see ECPDS Destination Authorization.
At a glance
- The plugin is read-only (
watch,replay). Thenotifyendpoint is never gated by ECPDS. - The plugin fails closed: it will never accidentally allow a request. The
status code distinguishes where the problem is.
503 Service Unavailablemeans the ECPDS check could not reach a verdict (an upstream / partial-outage problem); investigate ECPDS and the network.500 Internal Server Errormeans the plugin itself hit a server-side bug or a misconfiguration on Aviso’s side (missingAuthSettings, no checker registered, an unexpected plugin error); investigate Aviso. The full mapping is in the response codes table below. - The plugin does not retry. A
503is the signal to investigate ECPDS; a500is the signal to investigate Aviso. - The cache lives in process memory. Restarting Aviso clears it. Replicas have independent caches.
- The default
partial_outage_policyisstrict: every configured ECPDS server must respond successfully or the call fails with 503. A single ECPDS server going away takes the whole plugin down. This is intentional. The destination list itself is the union of every server’s response under both policies; the choice is purely about how tolerant we are of per-server failures.
Response codes the plugin emits
200 Allowed
The destination is in the user’s ECPDS allow-list.
Tracing event: auth.ecpds.check.allowed
403 Authorization denied
The destination is not in the user’s allow-list
(reason=DestinationNotInList), or the request omitted the configured
match_key field (reason=MatchKeyMissing).
Tracing event: auth.ecpds.check.denied
503 ECPDS or network failure
The combined ECPDS responses could not satisfy the active
partial_outage_policy. Check ECPDS availability, network connectivity and
service-account credentials. The event’s fetch_outcome field helps narrow
down the cause.
Tracing event: auth.ecpds.check.unavailable
500 Aviso error
A server-side bug or local misconfiguration prevented the check. Possible
causes include missing AuthSettings or EcpdsChecker in app_data, or an
unexpected plugin error.
Tracing event: auth.ecpds.check.error
Symptom and first checks
Start with one affected request and find its plugin event in the logs. An HTTP
403 or 503 alone does not show that ECPDS caused it. Use the same time window
and Aviso replica when comparing logs with metrics. The username values show
who was affected, not what caused the problem.
Watch or replay returns 503
event_name=auth.ecpds.check.unavailable confirms that the plugin could not
get a usable destination list under the configured partial_outage_policy.
- Find
event_name=auth.ecpds.fetch.failedfor that user and time. Readserveranderrorto identify the failing ECPDS server and its error. - Check that server from the Aviso host, using the configured service account
and the affected username as the lookup ID. For connection or timeout
errors, check the URL, DNS and connectivity. For HTTP 401 or 403, verify the
service account’s credentials and access. For other HTTP errors, inspect
the status: check the URL for 404, throttling for 429, and upstream logs for
5xx. For an invalid response, inspect the returned body and its
successvalue rather than assuming the API changed. - Check
partial_outage_policy:strictneeds every server to succeed;any_successneeds at least one. Use the per-server failure logs to see which servers need attention.
Metrics:
aviso_ecpds_access_decisions_total{outcome="unavailable"} counts requests
that the plugin rejected with 503. aviso_ecpds_fetch_total, grouped by
outcome, counts fetch attempts across the configured servers, not individual
server calls. Labels include unreachable, http_401, http_403, http_4xx,
http_5xx and invalid_response. The request log uses fetch_outcome with
values such as Unreachable or Unauthorized; the fetch failure log uses
error, not outcome or fetch_outcome.
Failed lookups are not cached, but concurrent requests for the same user can
share one fetch. The request and fetch counters need not rise together. Under
any_success, a failure label on the fetch metric can also accompany a usable
list, so it does not by itself mean a request returned 503.
Watch or replay returns 403
event_name=auth.ecpds.check.denied with reason=DestinationNotInList means
the requested destination was absent from the list Aviso used for that user.
That list may have come from cache, not a new ECPDS call. If the reason is
MatchKeyMissing, use the next section instead.
- Check the event’s
usernameandevent_typeagainst the intended user and schema. Compare the request’s destination with the configuredmatch_keyand any schema rules that change its value before the check. - Query the configured ECPDS servers using Aviso’s service account and that
username as the lookup ID. Check that the destination record has
active: trueand a string value intarget_field. Debug eventsauth.ecpds.fetch.skipped_inactiveandauth.ecpds.fetch.skipped_recordidentify servers whose records were excluded. Withany_success, also checkauth.ecpds.fetch.failed: a failed server’s destinations are absent. - Read
cache_outcomeon the denial.hitmeans Aviso reused the user’s cached list. If ECPDS access was recently changed, retry aftercache_ttl_secondsexpires. Check the same replica, since each has its own cache.
Metric:
aviso_ecpds_access_decisions_total{outcome="deny_destination"} counts denied
requests, including repeated checks against a cached list. Those cache hits
do not increase aviso_ecpds_fetch_total. This does not tell you how many
users are affected; use the denial logs for that.
Logs show MatchKeyMissing
event_name=auth.ecpds.check.denied with reason=MatchKeyMissing means the
configured match_key was absent from the processed request identifiers.
The plugin returns 403 before consulting the cache or ECPDS.
- Use
event_typeto identify the schema. Compareecpds.match_keywith its identifier field name and the actual request body. - Check the deployed version and the configuration loaded at startup.
Startup validation requires the match key to exist in the schema with
required: true; normal request validation rejects an omitted required field before the plugin runs. This event alone does not explain how the key went missing. - If those settings match, keep the request ID and a redacted request example for an Aviso bug report. Investigate how request processing passed identifiers without the key to the checker, rather than changing ECPDS permissions.
Metric:
aviso_ecpds_access_decisions_total{outcome="deny_match_key_missing"} counts
these denials. The denial event has cache_outcome="none"; this path does not
increase the cache or fetch counters.
Requests succeed, but no ECPDS checks appear
An absent auth.ecpds.check.allowed event does not prove the plugin is off.
Admin requests bypass the destination check, and log filters can hide events.
- Confirm that you are looking at new watch or replay requests for the
intended
event_type, not an already-open stream ornotifytraffic. Check the logs and metrics for the replica handling those requests. - Check
aviso_ecpds_access_decisions_total{outcome="admin_bypass"}. An increase explains why there are no allow events for admin requests. The matchingauth.ecpds.admin.bypassevent is debug-level; the normalauth.ecpds.check.allowedevent is info-level. Check the logging filters. - For a non-admin request, verify that the deployed schema’s
authblock containsplugins: ["ecpds"]andrequired: true. Startup rejects an ECPDS plugin reference if the binary lacks theecpdsfeature or the schema hasauth.required: false; those are not silent bypass settings.
Metrics: compare changes in aviso_ecpds_access_decisions_total by
outcome, not just allow. If all aviso_ecpds_* series are missing, verify
that metrics are enabled and the scrape reaches the right replica before
checking whether the binary was built with --features ecpds.
Starting a watch or replay is slow
event_name=auth.ecpds.cache.miss means a request could not use a cached list.
It may have fetched from ECPDS or waited for another request’s fetch. A miss
alone does not prove ECPDS caused the delay.
- Check
aviso_http_request_duration_secondsforroute="/api/v1/watch"orroute="/api/v1/replay". This measures time until response headers are ready, not how long the stream stays open. - For a slow request, inspect
cache_outcomeon itsauth.ecpds.check.*result event.hitmeans no upstream fetch was needed;miss_fetchedmeans this request fetched;miss_coalescedmeans it shared another request’s fetch. Checkauth.ecpds.fetch.failedfor timeout or connection errors. Cache hit/miss and fetch success events require debug logging. - Compare the rates of
aviso_ecpds_cache_misses_totalandaviso_ecpds_cache_hits_totalon that replica. Check recent restarts or traffic moving between replicas before tuning the cache. Comparecache_ttl_secondswith how often users reconnect, andaviso_ecpds_cache_sizewithmax_entries. The size gauge is an approximate count of cached usernames, not proof of eviction.
Use aviso_ecpds_fetch_total to check whether more misses also mean more
fetches; shared fetches count once. If slow requests are cache hits, continue
with the request’s other logs rather than assuming the cache needs resizing.
Tracing event reference
Every event uses the codebase’s standard structured shape (service_name,
service_version, event_name, plus event-specific fields). The list below
covers each event with a one-line meaning. Field-value details follow.
| Event | Level |
|---|---|
auth.ecpds.check.started | debug |
auth.ecpds.check.allowed | info |
auth.ecpds.check.denied | warn |
auth.ecpds.check.unavailable | warn |
auth.ecpds.check.error | error |
auth.ecpds.admin.bypass | debug |
auth.ecpds.cache.hit | debug |
auth.ecpds.cache.miss | debug |
auth.ecpds.fetch.succeeded | debug |
auth.ecpds.fetch.failed | warn |
auth.ecpds.fetch.skipped_inactive | debug |
auth.ecpds.fetch.skipped_record | debug |
auth.ecpds.check.started
The plugin started checking access for a request.
auth.ecpds.check.allowed
The plugin allowed the request.
auth.ecpds.check.denied
The plugin denied the request. See reason field.
auth.ecpds.check.error
An unexpected error in the plugin. See error_kind or error field.
auth.ecpds.admin.bypass
An admin user skipped the ECPDS check. Demoted from info because admin bypass is
configured behaviour, not an event SREs alert on; the
aviso_ecpds_access_decisions_total{outcome="admin_bypass"} Prometheus counter
still records every occurrence unconditionally.
auth.ecpds.cache.hit
The destination list came from cache.
auth.ecpds.cache.miss
The destination list was not in cache; a fetch was triggered.
auth.ecpds.fetch.succeeded
A fetch to one ECPDS server succeeded.
auth.ecpds.fetch.failed
A fetch to one ECPDS server failed. See error field.
auth.ecpds.fetch.skipped_inactive
One or more ECPDS records returned by a single server had active != true
(false, missing, or not a boolean) and got dropped from the user’s allow-list.
Carries server_index, server, username, skipped, total. Demoted from
info because every ECPDS fetch routinely returns inactive records and the skip
behaviour is the documented contract; flip to debug only when investigating a
denied user whose expected destination appears in this skip count.
auth.ecpds.fetch.skipped_record
One or more ECPDS records returned by a single server were active but missing
the configured target_field and got dropped. Carries server_index, server,
username, target_field, skipped, total so on-call can pinpoint which
ECPDS server is producing the malformed records. Demoted from info on the same
grounds as skipped_inactive.
Common fields
Most events carry event_type (the schema name) and username (the JWT
subject). Per-server events (auth.ecpds.fetch.succeeded, .failed, and
.skipped_record) also carry server_index (zero-based) and server (the
parsed URL).
Field value reference
Some events carry a typed enum field. The values you will see in logs are listed below. They are spelled exactly as shown.
reason(onauth.ecpds.check.denied):DestinationNotInList: the user is not entitled to the requested destination.MatchKeyMissing: the request body did not include the configured match-key field.
fetch_outcome(onauth.ecpds.check.unavailable):Unauthorized,Forbidden: an ECPDS server returned 401 or 403.ClientError: an ECPDS server returned a 4xx other than 401 or 403 (commonly 404 for a misconfigured base URL or 429 for throttling).ServerError: an ECPDS server returned 5xx.InvalidResponse: an ECPDS server returned a body the parser could not read.Unreachable: network or timeout failure.
cache_outcome(on everyauth.ecpds.check.*event:.allowed,.denied,.unavailable,.error):hit: served from cache.miss_coalesced: the cache was empty for this key but a concurrent caller’s fetch was in flight; this request waited on it.miss_fetched: this request ran the upstream fetch itself. The merged per-server result of that fetch is recorded as theoutcomelabel on theaviso_ecpds_fetch_totalmetric, and on theauth.ecpds.check.unavailableevent also asfetch_outcome(see above). It is intentionally NOT inlined intocache_outcomeso log filters keyed oncache_outcome:miss_fetchedstay stable as newFetchOutcomevariants are added.none: cache lookup was deliberately skipped because the request fell at theMatchKeyMissingdeny path before any cache call ran. Only appears onauth.ecpds.check.deniedevents alongsidereason=MatchKeyMissing.
How to confirm “config error vs. upstream outage”
-
Is the ECPDS plugin even compiled in? Check
/metricsforaviso_ecpds_*series. The unlabelled counters and gauge plus the pre-initialised label values onaviso_ecpds_access_decisions_totalandaviso_ecpds_fetch_totalregister at process startup whenever the binary is built with--features ecpds, regardless of whether anecpds:config block exists. If the series are absent, the binary does not have the feature, OR the metrics endpoint itself is disabled (metrics.enabled: falsein your config). If the series exist butaviso_ecpds_access_decisions_total{outcome="allow"}plusoutcome="deny_*"are all flat at zero under load, the plugin is compiled in but no stream actually opts in viaplugins: ["ecpds"]. -
Are the configured server URLs reachable from this Aviso host? Run this from the same host as Aviso:
curl -i -u "<service-username>:<service-password>" \ "https://<your-ecpds-host>/ecpds/v1/destination/list?id=<some-test-username>"200with a JSONdestinationList: ECPDS is up and credentials are valid. Problem is on the Aviso side.401or403: service-account credentials are wrong (rotated, revoked, typoed).5xxor hang: ECPDS itself is broken.- DNS error or connection refused: network-level issue.
-
Is one specific user being denied while others succeed? Run the curl above with that user’s id and compare with the destination they tried to read.
When an ECPDS server is unavailable
With the default partial_outage_policy: strict, every configured ECPDS
server must respond successfully when Aviso fetches a user’s destination
list. If one fails, requests needing a fresh list receive HTTP 503. Requests
that can use a cached list can still be checked against that list.
With partial_outage_policy: any_success, Aviso can use the lists from the
servers that respond successfully. A destination known only to a failed
server may be missing, so a user could be denied access they would normally
have. If no server succeeds, the lookup still fails with HTTP 503.
Both policies combine the destination lists returned by ECPDS. The choice is whether a lookup requires all servers or at least one. See Partial outage policy for more details.