Core Features
WebSocket Capture
Reliable instruments window.WebSocket to record one summary row per connection, lifecycle, message and byte counts, reconnect storms, and forwards abnormal closes and post-open errors into the errors pipeline.
How it works#
On init(), Reliable wraps window.WebSocket. Every connection created from that point is observed: the open timestamp, every message event (counted, not stored), every error, the close code, and outbound send() calls.
Counters live in memory for the lifetime of the socket. When the connection closes, cleanly or otherwise, a single POST /websocket event is emitted with the full summary. One row per connection, never one per message.
Why one row per connection?
Pre-upgrade handshakes#
A WebSocket starts life as an HTTP GET with an Upgrade: websocket header. If that fails with a 4xx or 5xx (auth, CORS, server unavailable) the upgrade never happens, and that request is already captured by the network module as a regular failed HTTP call.
The WebSocket module only takes over after the 101 Switching Protocols response. So a failed connection attempt shows up in Network, while a connection that opened and then dropped shows up in WebSockets.
What gets sent#
| Name | Type | Default | Description |
|---|---|---|---|
| url | string | — | Full WebSocket URL (ws:// or wss://). Sensitive query params like tokens or keys are scrubbed. |
| url_template | string | — | Normalized path, numeric, UUID, and long-hex segments collapse to :id / :uuid / :hex so /chat/users/42 and /chat/users/99 group together. |
| protocols | string[] | — | Sub-protocols negotiated at open (the second argument to new WebSocket(url, protocols)). |
| opened_at | string | — | ISO timestamp when the connection was created. |
| closed_at | string | — | ISO timestamp when the connection closed. |
| duration_ms | number | — | Lifetime of the connection in milliseconds. |
| close_code | number | — | WebSocket close code (1000 normal, 1001 going away, 1006 abnormal, 1011-1014 server error, 4000-4999 app-defined). |
| close_reason | string | — | Optional reason string the peer included with the close frame. |
| had_error | boolean | — | Server-computed verdict: true if any error event fired during the lifetime OR the close code is anything other than 1000/1001. |
| error_count | number | — | How many error events fired during the connection's lifetime. |
| messages_sent | number | — | Total messages your code sent over the connection. |
| messages_received | number | — | Total messages received from the peer. |
| bytes_sent | number | — | Sum of outbound payload sizes (strings are measured as UTF-8 bytes; Blob/ArrayBuffer use their byte length). |
| bytes_received | number | — | Sum of inbound payload sizes. |
| reconnect_count | number | — | Number of prior opens to the same url_template within a 30-second sliding window. A value > 3 signals broken backoff. |
| path | string | — | The page route the user was on when the connection opened. |
Errors flow into the errors pipeline#
A WebSocket summary on its own doesn't page anyone, but the failures inside it do. Three classes of events are forwarded to the standard errors pipeline so they participate in incident grouping, alerts, and paging:
- Post-open
errorevents: anything the WebSocket spec surfaces via theerrorhandler. - Abnormal closes: close codes outside
1000(normal) and1001(navigation). This includes1006(the socket died without a close frame, usually a network drop) and1011-1014(server-side errors). send()on a closed socket: a programming bug worth surfacing, even if the underlying call already throws.
Each forwarded error is tagged with source: "websocket", the URL, and the connection UUID, so you can filter the Errors view to just WebSocket failures, or pivot from a WebSocket summary row into the matching error in one click.
Reconnect storm detection#
When a new connection opens to a URL template that's already opened recently, the SDK stamps the running count onto the new row. Backoff is healthy when this stays at 0-2; values above 3 within the 30-second window mean the reconnect loop is hot and the dashboard will start flagging them.
// Healthy lifecycle — single open, clean close
{ reconnect_count: 0, close_code: 1000, duration_ms: 480_000 }
// Suspect — repeated opens to the same endpoint inside 30s
{ reconnect_count: 5, close_code: 1006, duration_ms: 1_200 }Pattern anomaly detection#
Alongside the summary row, the SDK sends a small per-connection sketch when the socket closes. Every message is reduced to a fingerprint: a hash of its structure (the sorted top-level JSON keys plus the value of a type-like field such as type or event, the first 16 bytes and length of a binary frame, or text with numbers and ids stripped). Payloads are never sent, and a fingerprint cannot be turned back into the message it came from. The sketch holds per-fingerprint message counts, payload-size and inter-message-gap percentiles, and which sent message types were answered by which received ones, with their round-trip times.
The backend learns a baseline per project from these sketches. Scoring starts once 1,000 sessions have been seen (the WebSockets page shows a calibration bar until then), and a fingerprint or request/response pair is only scored once it has appeared in 20 sessions. If more than 80% of traffic falls into the overflow bucket for rare shapes (encrypted or random-key payloads), anomaly detection switches itself off for the project and the page says so; lifecycle monitoring keeps working.
Signals
- Frequency: a message type was sent or received far more or less often than usual.
- Size: payloads of a message type were unusually large or small.
- Delay: the gap between consecutive messages of one type was unusual. This is not a response time.
- Adjacency: how often a sent message type was answered by a given received type was unusual.
- Round-trip: the time from a sent message to the response that answered it was unusual. Sessions flagged before this signal existed do not show it.
Reading the scores
Scores are unitless, and they are not standard deviations. Size, delay and round-trip compare the session's median with the baseline median, divided by the distance from the baseline median to its 99th percentile, so 1.0 means "as far from typical as the p99". Frequency and adjacency use a rough count-based distance on about the same scale. Each fingerprint or pair takes its worst signal, those are combined into one score per session, and a session is flagged when that score reaches 3.0.
Expanding a flagged session lists its top contributors: the direction (sent or received), what the fingerprint reveals about the message (JSON object, text, a binary frame and its size), a short hash you can copy to match the same shape in other sessions, the worst signal and its score. A request/response pair is shown as sent → received.
Toggling it#
WebSocket capture is enabled by default. Disable it via the captureWebSockets flag if you don't use WebSockets or want to gate it explicitly:
init({
publicKey: 'pk_live_rl_...',
captureWebSockets: false, // default: true
});On the server side, WebSocket data is accepted only while Network Requests capture is on in the project's settings. Turning that off stops both network and WebSocket ingest, and the WebSockets page says so.
What's NOT captured#
- Per-message payloads: only counts, byte sizes and the structural fingerprints described above. We never store frame contents.
- Sockets opened by browser extensions or other libraries that bypass
window.WebSocket: the wrapper applies atinit(); anything created before that, or that holds a reference to the original constructor, escapes capture. - Server-Sent Events (SSE) and EventSource: those go through fetch and are covered by the network module instead.
Tip
beforeSend(event) {
if (event.path === '/websocket' && event.url_template?.endsWith('/heartbeat')) {
return null;
}
return event;
}