Ask Ela
Streams an answer grounded only in the caller’s own library as text/event-stream: one retrieval, then delta events, then one done or error. Zero strong matches returns “I couldn’t find this in your notes” without a model call. Citations [n] refer to retrieval.hits[n-1]; citations to anything not retrieved are dropped. 404 when Ask Ela is not rolled out to the caller, 403 PLAN_REQUIRED without Ela Pro, 503 ASK_UNAVAILABLE when the server has no model key. Prompts and completions are never logged or stored.
Authorizations
Static bearer, WorkOS JWT, or open local when AUTH_REQUIRED is off and no static tokens are set. The token selects the vault; paths do not include vault_id.
Body
1 - 1000Response
text/event-stream. The body is a sequence of SSE frames, not a single JSON object.
Framed SSE text, not one JSON value. Each message is event: <name> then data: <json> then a blank line. Lines starting with : are keep-alives. Order: one retrieval (hits, coverage, ranker_version), zero or more delta (text), then exactly one done (answer, refused, citations, model) or error (code, message). error.code is ASK_MODEL_UNAVAILABLE, ASK_UPSTREAM_ERROR, ASK_TIMEOUT, ASK_INCOMPLETE, or ASK_INTERNAL. ASK_INCOMPLETE means the model stopped early (max_output_tokens or a content filter); that stream is not a finished answer. Read the body as text/event-stream. TypeScript clients pass parseAs "stream" (see askNotesStream). Do not JSON.parse the whole payload.
"event: retrieval\ndata: {\"hits\":[],\"coverage\":{\"notes_searched\":1,\"strong_match_count\":0,\"excluded_folders\":[\"private/\"],\"layer_sequences\":{}},\"ranker_version\":\"lexical-rrf-v1\"}\n\nevent: delta\ndata: {\"text\":\"See \"}\n\nevent: done\ndata: {\"answer\":\"See [1].\",\"refused\":false,\"citations\":[{\"slug\":\"notes/paddle\",\"line_start\":3,\"line_end\":7}],\"model\":\"gpt-6-luna\"}\n\nevent: error\ndata: {\"code\":\"ASK_INCOMPLETE\",\"message\":\"Ask Ela's answer was cut off before it finished. Try again.\"}\n"