11. Unresponsive endpoints & partial results
11.1 Primary carrier: meta.federation.endpoints[]
Every node that was in scope MUST appear in meta.federation.endpoints[] with a status (and an error where relevant), whether or not it contributed rows (N16). A registry member that was not in scope SHOULD appear as well, as excluded or not-localized, so that a client sees the whole federation and not a list silently truncated to the nodes that were asked; a gateway that omits them MUST NOT emit them under any other status. Unresponsive nodes MUST NOT contribute rows. The status set is:
| status | Meaning |
|---|---|
|
Queried and responded. |
|
Known node, not reachable. |
|
Reachable but did not respond within the gateway timeout. |
|
Reached and answered, but with a failure instead of a result set: an HTTP error status, a malformed body, or a response the gateway could not use. The node’s own error is carried in |
|
The cross-reference service returned no local |
|
Consent did not permit this node’s inclusion. Set either by an optional Step-1 consent pre-filter (N27a, §13.2.1) or by the node itself refusing on consent grounds (N27); deployments without a consent service use it only in the latter sense. |
|
Excluded by a decision about this node: a directive that did not name it, say a directed |
|
A registry member that localization did not return as a candidate for this query (§14). The node was never asked, and nothing decided about it: no directive, no operator policy, no consent decision. An undirected query produces this status routinely and it is not a failure: a localizer naming three of nine members leaves six |
The difference between the last two matters for audit. excluded records that something ruled the node out; not-localized records that nothing ruled it in. Neither means the node was asked and declined, which is consent-denied.
The three statuses that describe an attempted request are likewise distinct on purpose. offline means no connection; time-out means a connection but no answer in time; node-error means an answer that was a failure. They map to different HTTP statuses under the default strategy (§11.4) and call for different recoveries, so a gateway MUST NOT fold a node error into offline or report it as active with an error attached.
What "in scope" means. A node is in scope for a query if node selection put it there, whether localization named it or a directive did. A node reported excluded or not-localized was never in scope: one because a decision removed it before selection completed, the other because nothing selected it. Both are reported, so a client sees the whole federation picture and not a silently truncated list, but neither clears meta.federation.complete (§11.4, N37), because a query is not incomplete for failing to ask a node it never intended to ask. Every other status describes a node that was in scope and did not reach active, and so does clear the flag.
Which statuses fail the query. Under the all-or-nothing default of §11.4, an in-scope node not reaching active does more than clear a flag: it fails the query (504 or 424). That consequence does not extend to every in-scope status, and the distinction is normative:
| status | In scope? | Consequence |
|---|---|---|
|
yes |
Contributed; nothing to report. |
|
yes |
Fails the query under the default strategy, with |
|
yes |
Fails the query under the default strategy, with |
|
yes |
Clears |
|
yes |
Clears |
|
no |
Neither clears |
11.2 HTTP status mapping (aligned with Exchange-Routing)
The Tier maps errors as follows, aligning with the Exchange-Routing IG’s exception table, which separates intermediary-reported from destination-reported exceptions.
This table maps conditions that fail the request. Under the default all-or-nothing completion strategy (§11.4) a single in-scope node that was asked and did not answer does fail the request. Where a caller opted into best-effort with openEHR-federation-completeness: partial, the same condition is reported in meta.federation.endpoints[] and the request succeeds. The rows below that concern a single node’s failure therefore name the completion strategy they assume. Where the two could be read differently, §11.4 governs: it is the more specific treatment, N37 restates it in the normative requirement set, and a client can be written against it.
Two in-scope statuses are absent from this table and do not fail the request in either mode: not-resolved and consent-denied (§11.3), because a node that answered "not this patient" or "not to you" has answered.
| Condition | HTTP | Origin |
|---|---|---|
Patient/endpoint cannot be resolved to a destination |
404 |
intermediary (gateway) |
Node connection timed out or is unreachable (no response), under the default all-or-nothing strategy (§11.4) |
504 |
intermediary (gateway); the failing response still carries |
Node returned an error, under the default all-or-nothing strategy (§11.4) |
424 |
intermediary (gateway); the node is reported |
Node connection timed out (no response), where best-effort was opted into ( |
200 |
intermediary (gateway); the node is reported |
Gateway internal error |
500 |
intermediary (gateway) |
Node returned not found for the submitted path/resource |
404 |
node (passed through) |
Invalid request (e.g. bad AQL, un-reducible subject predicate, §7.1) |
400 |
gateway or node |
Node internal error |
500 |
node (passed through) |
|
409 |
intermediary (gateway) |
An ITS-REST area the gateway does not expose (N32) |
501 |
intermediary (gateway) |
A node’s internal error maps to 500 and passes through, matching the Exchange-Routing IG’s destination-reported column. The error is the destination’s, and 502 Bad Gateway would pin it on the intermediary. A gateway MAY separate the two in the per-endpoint error field, which carries per-node detail.
11.3 "Patient found nowhere" is not a 404 - and not a 424
If no node resolves the patient, so every endpoint is not-resolved, the correct response is HTTP 200 with an empty rows array and an explanatory meta, not a 404. A 404 is reserved for a gateway that cannot resolve a request to any destination at all. The split keeps "the patient legitimately has no record in this federation" apart from "the request was unroutable."
not-resolved is a legitimate empty answer, not a failed dependencyThe all-or-nothing default of §11.4 does not override this rule, although a naive reading of that default would break it.
The carve-out, normatively: a node reported This agrees with N6, which has always said |
A consent-denied node follows the same logic for the same reason. Consent is the node’s gate (N27), a refusal is a decision, not an outage, and a query MUST NOT fail because one node declined to release data it holds. Like not-resolved it clears meta.federation.complete, so the client is told coverage was incomplete, but it does not produce a 424. A federation in which one patient’s opt-out at one node failed every query about them would be unusable, and would leak that opt-out through the error.
11.4 All-or-nothing is the default completion strategy (normative)
A fan-out has to decide what to do when some nodes answer. This specification fixes that decision, because a client cannot interpret a result set without knowing which rule produced it.
-
The default strategy is all-or-nothing: if an in-scope node (§11.1) was asked and did not answer - status
offlineortime-out- or answered with an error - statusnode-error- the gateway MUST fail the query instead of returning the rows it did obtain. The status is504where the cause is a timeout or unreachability and424 Failed Dependencywhere a node returned an error. Where one fan-out has both kinds,504takes precedence, because an unanswered node is an unknown and a retry or apartialrequest is the client’s cheapest recovery, whereas a node error needs no retry to be read. The two in-scope statuses that represent an answer -not-resolvedandconsent-denied- do not fail the query; see the carve-out in §11.3 and the table in §11.1. (N37.)Why this is the default. A partial clinical answer that looks exactly like a complete one is a safety hazard, and the hazard is asymmetric. A clinician who does not know a node was missing may reasonably conclude the data does not exist, that there is no allergy, no prior admission, no anticoagulant, and act on that absence. The flag of [completeness-flag] makes incompleteness detectable but not detected: it relies on every client, in every rendering path, checking a boolean nothing forces it to read. Failing closed removes the dependence on that discipline: a client that asked for the federation’s answer and cannot be given it is told so.
The same reasoning makes localization fail closed (§14.1), so that a degraded federation is visible as degraded instead of appearing empty.
-
The gateway MUST make coverage machine-detectable, not merely inferable by scanning statuses.
meta.federationMUST carry acompleteboolean,trueonly when every in-scope node reached statusactive. A node reportednot-localized(§11.1) was never in scope for this query and does not clearmeta.federation.complete. Localization normally names only a subset of the federation; that is not an incomplete answer, and a flag that went false on nearly every undirected query would tell a client nothing.completeanswers "did every node I asked answer?", not "did I ask every node?" The flag is required in both modes, and under neither is it redundant with the HTTP status: on the default strategyfalseappears on a failing response, and also on a200where the only in-scope statuses short ofactivewerenot-resolvedorconsent-denied(§11.3). A client MUST therefore read the flag instead of inferring coverage from the status code. For FHIR-facing consumers, incompleteness also surfaces as anOperationOutcomewarning (CP-12). (N37.) -
A failing response MUST still carry the diagnostic envelope. A
504or424under the default strategy MUST carrymeta.federation.endpoints[]with the per-node status of every in-scope node, andmeta.federation.complete: false. A client has to be able to see which node failed and why, so it can retry, route around it, raise it with the operator, or decide apartialanswer would do after all. That only works if the error body keeps the envelope instead of collapsing to a bare status line. Of the requirements in this section, this one is the most likely to be dropped in implementation, and a gateway that drops it turns a diagnosable outage into an opaque one. (N37, CP-30.) -
A deployment MAY offer best-effort as an opt-in mode, in which the gateway returns the rows it obtained from the nodes that answered within the timeout budget (§11.5), reports every other node in
meta.federation.endpoints[]with its status, setsmeta.federation.complete: false, and responds200. If offered it MUST be selected per request - theopenEHR-federation-completeness: partialrequest header - and MUST be declared inOPTIONS {base}/(§7a.2). (N37.)Why the mode still exists. A flagged partial answer can be more useful than no answer. A clinician who can see four of five records, and can see that the fifth system is down, is better served than one who gets an error, provided the missing coverage is shown to them and not merely present in the payload. Whether that trade-off is acceptable depends on the deployment and its clients, so the mode remains available and is opt-in: a caller that asks for
partialhas declared that it will handle a partial answer.Compared with earlier releases only the default changed: a client that sends no completeness header gets the all-or-nothing behaviour.
-
The
openEHR-federation-completenessheader carriesall, the default and now implicit, orpartial. A gateway MUST acceptallexplicitly even though it is the default, so a client can state its requirement without relying on the default. A gateway that does not offer best-effort MUST rejectpartialinstead of silently ignoring it. -
Both strategies apply to reads. Neither applies to writes: a write is routed to exactly one node (§12.4) and either succeeds or fails there. There is no partial write in v1.
-
A node that returns an error, as opposed to not answering, is treated the same way for completeness purposes. It contributes no rows, appears as
node-errorwith the node’s error, and clearsmeta.federation.complete, failing the query with424under the default or being reported underpartial.
11.5 Timeouts (normative)
Timeouts are in scope. Leaving them to implementers would make the completeness semantics of §11.4 untestable, and a client cannot set its own deadline against a gateway whose behaviour it cannot predict. This specification fixes the decision structure and its visibility; the numbers are a deployment matter.
-
A gateway MUST operate with (a) a per-node timeout and (b) an overall query budget, and MUST apply both. When a node exceeds the per-node timeout, the gateway MUST abandon that node, mark it
time-outinmeta.federation.endpoints[], and continue with the rest. When the overall budget expires, the gateway MUST stop waiting for all outstanding nodes, mark eachtime-out, and return what it has under §11.4. (N38.) -
The values in force MUST be discoverable:
OPTIONS {base}/reportstimeout.per_node_msandtimeout.overall_ms(§7a.2), and every response MUST report, per endpoint, the elapsed time the gateway observed (§9.5). (N38.) -
A client MAY request a shorter budget with the standard
Prefer: wait=<seconds>header. The gateway MUST honour a shorter client budget and MUST NOT extend beyond its own overall budget on a client’s request, since a client cannot lengthen the gateway’s deadline. The effective budget MUST be reported inmeta.federation.timeout. -
The gateway MUST NOT let a node’s timeout cascade: abandoning a node MUST NOT abort in-flight requests to other nodes.
-
What a client can rely on. A conformant gateway answers within its declared overall budget, plus combining time. A
200means every node it asked that holds the patient answered, unless the client opted into best-effort, in which case the answer may be partial andmeta.federation.completesays so.meta.federation.completeandmeta.federation.endpoints[]identify which nodes are missing and why, including on a504/424(§11.4). Atime-outornode-errornode means unknown, never no data, and a client MUST NOT read such an endpoint as an empty result; the per-endpoint status set exists so that these cases stay distinguishable.A client MUST NOT assume that a
200withcomplete: falsecannot happen. It happens whenevernot-resolvedorconsent-deniedcleared the flag without failing the query (§11.3). In that case the gateway is reporting incomplete coverage as §11.3 requires, and has not violated the default. -
A node that answers after being abandoned MUST be ignored; the gateway MUST NOT emit rows outside the response it already returned. (An async request, §11.7, is the supported way to wait longer.)
11.6 Ordering, LIMIT, and aggregation across nodes (normative)
ORDER BY, LIMIT/OFFSET and aggregate functions each execute inside a node (N9), so their results cannot be taken at face value across a fan-out. The three cases have three different answers.
11.6.1 ORDER BY + LIMIT - re-executed at the Tier
A federated LIMIT n cannot be pushed down as LIMIT n and then concatenated. Taking the first n rows of each node’s locally ordered set and concatenating them yields a set that is neither globally ordered nor the global top n.
-
For a query with
ORDER BYandLIMIT nand noOFFSET, the gateway MUST dispatchLIMIT nto each node, then apply theORDER BYacross the merged rows, and finally applyLIMIT nto the merged, re-ordered set. Dispatchingnper node makes this correct, because under a total order the global topnis necessarily contained in the union of the per-node topn. (N39.) -
The gateway MUST apply
ORDER BYat the Tier whenever more than one node contributed rows, even if it also pushed it down (N13). Pushing down is an optimisation; the merge at the Tier produces the true order. -
Ordering MUST be deterministic. Where the
ORDER BYkeys tie, the gateway MUST break the tie on a stable secondary key (RECOMMENDED:endpoint_id, then uid), so repeating a query returns rows in the same order.
11.6.2 OFFSET - not correct across nodes
-
LIMIT n OFFSET kwithk > 0cannot be made correct by the same trick, because the global rowsk…k+nare not contained in the per-node rowsk…k+n. A gateway MUST NOT silently pushOFFSETdown and present the result as a correct global page. It MUST either-
reject the query with
400and an error stating that offset-based paging is not supported across a fan-out; or -
compute the page correctly by retrieving
k + nrows per node, merging, ordering and slicing - permitted only where the gateway can boundk + n(it MUST reject when it cannot); or -
serve the page from a materialised, ordered result it is holding for that query (§11.6.4).
Whichever it does, it MUST declare which in
OPTIONS {base}/. (N39.) -
-
The rule above is the concrete form of the "cross-node pagination is not guaranteed" non-goal (§2.3). Not guaranteed means the gateway may not be silently wrong: it must refuse or be correct, never approximate.
11.6.3 Aggregates - the query is wrong, not the gateway
COUNT, SUM, AVG, MIN, MAX fanned out and concatenated produce one row per node instead of a single aggregate. Summing them is right for COUNT/SUM, wrong for AVG, and right by accident for MIN/MAX only if the gateway knows to re-apply the function instead of concatenating.
-
Per N14 an undirected aggregate MUST be blocked. The gateway MUST reject it with
400and an error saying why, namely that the aggregate cannot be computed correctly across the fan-out, and SHOULD suggest the two supported alternatives: pin the query to one node (§8), or select the underlying rows and aggregate in the application. Returning per-node aggregate rows as if they were the answer is forbidden, because it is the only failure mode here that a client cannot detect. (N39.) -
A directed single-node aggregate is permitted and dispatched unchanged (N14).
-
A gateway MAY also support decomposable aggregates over multiple nodes:
COUNTandSUMby summing the per-node results,MIN/MAXby re-applying across them. The result must be exactly correct, and the aggregate must not be combined withDISTINCT,GROUP BYon a dimension that spans nodes, or de-duplication (§10), any of which breaks decomposability.AVGMUST NOT be decomposed this way unless the gateway also retrieves the per-node counts. A gateway supporting this MUST declare which functions inOPTIONS {base}/. A gateway that does not support it falls under the rule above, unchanged. -
GROUP BYfollows the same logic. Groups from different nodes with the same key MUST be merged at the Tier, or the query MUST be rejected, because emitting the same group key twice is a wrong answer.
11.6.4 Materialised result sets (optional)
A gateway MAY hold an ordered, merged result set for a bounded period and serve OFFSET-based pages from it, as a cursor. If it does, meta.federation MUST carry the cursor handle and its expiry, and the gateway MUST NOT re-run the fan-out mid-cursor, because a page served from a re-run is not a page of the same result. The facility is OPTIONAL, and a deployment can use it to offer honest cross-node pagination today. A general solution is deferred (§18).
11.7 Asynchronous long-running queries (optional)
For long fan-outs, the Tier MAY support the Exchange-Routing async pattern: the client sends Prefer: respond-async, the Tier replies 202 Accepted with a Content-Location polling URL, and the client polls until 200 OK with the result. The pattern is OPTIONAL and does not change the result envelope (§9). Async is the supported way to exceed the synchronous budget of §11.5. It does not exempt the gateway from that budget on ordinary requests.