GoodMemGoodMem
How-To Guides

REST timeouts and cancellation

Set a request budget and understand cancellation of REST responses and retrieval work.

REST timeouts and cancellation

Send the optional GoodMem-Timeout-Ms header on any GoodMem REST resource request to limit how long its internal gRPC calls may wait. This applies to reads, writes, and streaming retrieval through the shared REST-to-gRPC bridge.

curl --max-time 35 \
  -H "x-api-key: $GOODMEM_API_KEY" \
  -H 'GoodMem-Timeout-Ms: 30000' \
  "$GOODMEM_URL/v1/spaces"

The header accepts one decimal integer from 1 to 86400000 (one day), measured in milliseconds. Duplicate headers, zero, negative values, fractions, and values outside this range receive HTTP 400. Header names are case-insensitive.

The budget starts when the REST request enters the server's servlet processing, so time spent reading or converting the request consumes the budget. Every gRPC call made for that request shares the same absolute deadline; subsequent calls do not restart the timer. An existing shorter gRPC deadline still takes precedence. The budget bounds waiting for gRPC, rather than every part of HTTP body upload or response transmission.

Omitting the header adds no gRPC deadline. Existing server, proxy, provider, and HTTP client timeouts still apply. The header does not configure the independently mounted inference proxy (/v1/proxy/...), hosted MCP (/mcp), health, metrics, or console routes.

Deadline responses

If a gRPC deadline expires before the response starts, GoodMem returns HTTP 504 with its usual JSON error body. Once a streaming response has started, its HTTP status cannot be changed. SSE sends a terminal error event when the connection remains writable. NDJSON can provide the same explicit outcome using the opt-in header described below. Legacy NDJSON cannot reliably distinguish complete results from a truncated response merely by observing EOF.

A local client timeout, such as curl --max-time or an aborted browser fetch, controls the client's wait. It does not communicate an execution budget to the server. Set the header as well when you want the server's gRPC calls to have a deadline. Allow additional client-side time to receive the server's timeout response.

Disconnects and operation outcomes

When Jetty reports a connection failure, servlet async timeout, or failed streaming write, the shared bridge cancels outstanding gRPC calls for that request. Network and intermediary behavior can delay disconnect detection, especially while the server is not writing a response. A silent provider call can run until completion before that disconnect is discovered. Always send a timeout header when you need a bounded wait and timely cancellation of cancellable read work; a client abort alone does not provide that guarantee.

RetrieveMemory and ping probes cooperatively stop their read work when their gRPC context is cancelled or their deadline expires. Streaming ping stops scheduling further probes. The bridge does not forcibly interrupt every service operation.

An accepted mutation can commit after a timeout or disconnect. A failed or missing response describes response delivery; it does not prove that a create, update, or delete was rolled back. An immediate GET returning 404 also does not prove the original mutation will never commit: it may still be running.

For create endpoints that accept a caller-supplied resource ID, choose that ID before the first attempt and reuse the same ID on a retry. Read and verify the resulting resource against the intended request. Reusing an ID prevents creating a second resource under a new ID; it is not a general guarantee that every side effect of a retried request is idempotent. HTTP 409 alone is not proof that the first attempt committed, because it can represent another conflict, such as a duplicate name. For other mutations, resolve the uncertain outcome according to the operation's semantics before retrying.

Detecting stream completion

For NDJSON retrieval or ping streams, send GoodMem-Stream-Terminal: true and verify the same header is echoed in the streaming response. This explicitly opts into transport outcome records in addition to the existing domain event records. Clients that need to distinguish complete results from truncated results must opt in; older servers without the echoed header do not provide this guarantee. The response ends with exactly one of these records when delivery succeeds:

{"type":"completed"}
{"type":"error","error":"Request timeout","status":504,"timestamp":1790467200000}

Handle a record's type before decoding a domain event. The error fields use the ordinary REST error contract; status describes the terminal failure, while the HTTP response itself remains 200 once streaming has started. A result-set END closes only that result set: an LLM reply may still follow. Only the request-level completed record establishes successful stream completion.

For SSE, the existing success frame remains unchanged:

event: close
data: {"type":"completed"}

A failure after streaming starts uses event: error with a type: "error" payload and the same error, status, and timestamp fields as above. The event: close frame remains exclusively a success marker, preserving existing SSE consumers. A disconnect can prevent any terminal record from arriving. In either format, EOF without a successful terminal record means the request's result is incomplete or its delivery is uncertain, even if some useful domain events were received. Legacy NDJSON remains unchanged unless the terminal header is enabled.