> ## Documentation Index
> Fetch the complete documentation index at: https://docs.blackbox.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmark Logs & Status

> Poll a benchmark run's progress, read its orchestrator logs, or stream them live over SSE. Logs come from memory while the run is alive and from durable storage after it ends.

Track a benchmark run end-to-end: a lightweight **status** poll, a buffered **logs** snapshot, and a live **SSE stream**. Logs are served **live from memory while the run is alive, and from the database after it ends** (or after an unexpected sandbox/process death), so they remain available post-completion.

## Authentication

All endpoints require a BLACKBOX API key as a Bearer token and are owner-scoped. See [Authentication](/api-reference/v1/authentication).

## Status — `GET …/runs/{id}/status`

A cheap progress poll.

<ResponseField name="status" type="string">`queued` | `preparing` | `running` | `completed` | `failed` | `cancelled`.</ResponseField>
<ResponseField name="progress" type="number">0–100.</ResponseField>
<ResponseField name="completedTasks" type="number">Tasks finished so far.</ResponseField>
<ResponseField name="totalTasks" type="number">Tasks in the run.</ResponseField>
<ResponseField name="resolvedTasks" type="number">Tasks that passed.</ResponseField>
<ResponseField name="resolvedRate" type="number | null">Fraction `0..1`, set on completion.</ResponseField>

```json status theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
{
  "benchmarkRunId": "b41ab8b2-…",
  "status": "running", "progress": 80,
  "completedTasks": 8, "totalTasks": 10, "resolvedTasks": 8,
  "resolvedRate": null, "error": null,
  "startedAt": "2026-06-13T22:47:00.782Z", "completedAt": null
}
```

## Logs snapshot — `GET …/runs/{id}/logs`

Returns the current orchestrator log lines.

<ResponseField name="source" type="string">
  `live` — served from the in-memory buffer (run still active). `persisted` — served from the durable DB snapshot (run ended and the buffer was cleared).
</ResponseField>

<ResponseField name="lineCount" type="number">Number of lines returned.</ResponseField>
<ResponseField name="logs" type="array">The log lines (each timestamped).</ResponseField>

```json logs theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
{
  "benchmarkRunId": "b41ab8b2-…",
  "status": "completed",
  "source": "persisted",
  "lineCount": 92,
  "logs": [
    "[2026-06-13T22:47:01Z] [prepare] creating builder sandbox...",
    "[2026-06-13T22:54:42Z] [task aime__65] graded numeric_match: answer=\"104\" gold=\"104\" → PASS",
    "[2026-06-13T22:54:46Z] [done] resolved 9/10 (resolved_rate=0.9000)"
  ]
}
```

## Stream — `GET …/runs/{id}/logs/stream`

A `text/event-stream` (SSE) of the logs.

* **While the run is alive:** replays the buffered lines, then pushes new lines as they arrive; closes with `event: end` when the run reaches a terminal state.
* **After the run ends:** replays the durable **DB** snapshot one-shot, then `event: end`. (This is why a completed run's stream is never empty.)

Each line is `data: <log line>\n\n`.

```bash stream theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
curl -N 'https://agent.blackbox.ai/api/v1/benchmarks/runs/RUN_ID/logs/stream' \
  -H 'Authorization: Bearer YOUR_API_KEY'
```

```text SSE theme={"theme":{"light":"github-light-default","dark":"github-dark-default"}}
data: [2026-06-13T22:47:01Z] [prepare] creating builder sandbox...

data: [2026-06-13T22:54:46Z] [done] resolved 9/10 (resolved_rate=0.9000)

event: end
data: {}
```

<Tip>
  Persistence is **incremental** — logs are mirrored to the database on a short throttle (plus a final flush when the run ends), so even if the sandbox or process dies mid-run, the DB holds a current snapshot. "Live from memory, dead from DB."
</Tip>

## Cancel — `POST …/runs/{id}/cancel`

Cancels an in-flight run; returns the new `status` (`cancelled`). Completed/failed runs are not affected.

## Tasks — `GET …/runs/{id}/tasks`

Per-instance results (`instanceId`, `status`, `reward`, `durationMs`, `error`, timestamps). For the aggregated score + traces in one call, prefer [Benchmark Results](/api-reference/v1/benchmark-results).

## Where logs live

Benchmark runs, task results, and logs are stored in a **dedicated benchmark database** (configured separately from the main app database). This keeps high-volume evaluation data isolated; runs and logs remain queryable by id after completion.

<CardGroup cols={2}>
  <Card title="Benchmark Results" icon="chart-simple" href="/api-reference/v1/benchmark-results">
    Score, percentage, per-task traces, tracking log.
  </Card>

  <Card title="Run Benchmarks" icon="play" href="/api-reference/v1/benchmarks">
    Launch a run (Claude / Codex / Grok, or your own router).
  </Card>
</CardGroup>
