Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Benchmarks

Perspective’s performance is tracked by a benchmark suite which lives in the repository. CI runs it on every tagged release and attaches the raw results to that GitHub release as Apache Arrow files — and the results are, naturally, explored in Perspective.

Results

Open the published results for the latest release, live in your browser:

BuildDashboardBy releaseRaw data
JavaScript (WebAssembly engine under Node.js)Benchmarks — JavaScriptHistorybenchmark-js.arrow
Python (native engine, driven over WebSocket)Benchmarks — PythonHistorybenchmark-python.arrow

Each file holds one row per timed iteration:

ColumnMeaning
benchmarkThe case, e.g. .view({group_by})
versionThe Perspective release the case ran against
version_idxRelease order; 0 is the build being released
real_time, cpu_time, user_time, system_timeMicroseconds
outliertrue when the iteration falls outside 1.5 × the interquartile range of its case

The dashboards are ordinary Perspective views over that file — mean real_time in milliseconds, excluding outliers, grouped by benchmark and version — so they can be re-pivoted, filtered to one case, or switched to another chart type in place.

Environment. Results are produced by GitHub Actions on an ubuntu-22.04 x86_64 hosted runner with Node.js 22 and Python 3.11, over the Superstore sample dataset. Hosted runners are shared, modest machines: read these numbers as a release-over-release trend on constant hardware, not as the ceiling for your own.

What is measured

The cross-platform suite (tools/bench/cross_platform_suite.mjs) defines the cases, which are written against the Client API and so can be pointed at any build of the engine:

AreaCases
Table constructiontable(arrow), table(csv), table(json), table(columns), and table(arrow, {limit})
Streamingtable.update(arrow), and table.update(arrow) with window columns active
Queriesview(), view({group_by}), view({group_by, aggregates: "median"}), view({expressions}), view({windows})
Joinsjoin()
Serializationto_arrow(), to_csv(), to_columns(), to_json()

A separate suite (charts_suite.mjs) measures chart rendering in a real browser.

Each case is run repeatedly against the current build and against previously published releases, so every number can be read relative to the versions before it. Results are written as Apache Arrow files under tools/bench/dist/.

Running it

From a built checkout of the repository:

cd tools/bench
pnpm run bench_js
pnpm run bench_python
pnpm run bench_charts

Reading results sensibly

  • Arrow is the fast path. Loading Arrow avoids parsing and type inference; CSV and JSON construction times measure the parser as much as the engine.
  • Column types matter more than row count. Numeric and datetime columns are fixed-width; string columns are dictionary-encoded and cost more to build and to group.
  • Updates scale with the delta. The cost of update() on a table with active views is driven by the size of the update and the number of groups it touches, not by the size of the table.
  • WebAssembly is single-threaded per worker; native is not. The Python, Node.js and Rust builds use a thread pool (perspective.set_num_cpus()), so native numbers are typically better than in-browser numbers for the same case.
  • Rendering is separate from querying. The data grid draws only the cells in view, so a grid over ten million rows renders in the same time as one over ten thousand; the query is what scales.

For guidance on sizing, see Visualizing millions of rows in the browser.