BasekickLabs

Grafana Integration

Connect Grafana to Arc Enterprise: install the data source, point it at a cluster endpoint, and scope what each dashboard audience can read with RBAC tokens.

Connect Arc Enterprise to Grafana for real-time dashboards, alerting, and ad-hoc analysis, using the Arc data source plugin.

The plugin is the same one the open-source edition uses. This page covers it end to end, and calls out the three places an Enterprise deployment differs: which endpoint to point at, how RBAC decides what a dashboard can read, and what query governance does to a runaway panel.

Overview

The Arc data source plugin gives Grafana:

  • Three wire protocols — Apache Arrow IPC (default), columnar MessagePack, and JSON
  • Native SQL — full DuckDB analytical SQL, not a query builder
  • Grafana macros — $__timeFilter, $__timeGroup, $__interval and friends
  • Timezone-aware bucketing — daily and weekly buckets align to the dashboard's local calendar
  • Template variables — dynamic dashboards with single- and multi-value filters
  • Alerting — full support for Grafana alert rules
  • Query splitting — long ranges are chunked and run in parallel, with automatic skipping where that would change results

Requirements

  • Grafana 12.3 or later
  • An Arc Enterprise endpoint reachable from the Grafana server
  • An Arc API token, scoped to what the dashboards should read

Installation

The plugin is not yet in the Grafana plugin catalog, so install it from a release or from source. Because it is unsigned, Grafana must be told to load it.

From a release

# Resolve the latest release tag, then download the matching plugin archive
LATEST=$(curl -s https://api.github.com/repos/basekick-labs/grafana-arc-datasource/releases/latest \
  | grep tag_name | cut -d '"' -f 4 | sed 's/v//')
wget https://github.com/basekick-labs/grafana-arc-datasource/releases/download/v${LATEST}/basekick-arc-datasource-${LATEST}.zip

# Extract into the Grafana plugins directory
unzip basekick-arc-datasource-${LATEST}.zip -d /var/lib/grafana/plugins/

systemctl restart grafana-server

Allow the unsigned plugin

Grafana refuses to load unsigned plugins by default. Add the plugin ID to your configuration:

# /etc/grafana/grafana.ini
[plugins]
allow_loading_unsigned_plugins = basekick-arc-datasource

Or, with Docker:

environment:
  GF_PLUGINS_ALLOW_LOADING_UNSIGNED_PLUGINS: basekick-arc-datasource

Without this, the plugin is downloaded and ignored, and the data source never appears in the list.

From source

git clone https://github.com/basekick-labs/grafana-arc-datasource
cd grafana-arc-datasource

npm install
npm run build     # frontend
mage -v           # backend (requires Go 1.26.6+)

cp -r dist /var/lib/grafana/plugins/basekick-arc-datasource
systemctl restart grafana-server

Configuration

1. Add the data source

  1. In Grafana, go to Connections → Data sources
  2. Click Add new data source
  3. Select Arc

2. Connection settings

SettingDescriptionRequiredDefault
URLArc API endpointYeshttp://localhost:8000
API KeyArc authentication tokenYes—
DatabaseDefault database for this data sourceNodefault
TimeoutQuery timeout in secondsNo30
ProtocolArrow, MessagePack, or JSONNoArrow
Max ConcurrencyParallel chunks within one split queryNo4
Max In FlightSimultaneous Arc requests for this data sourceNo32
Max Response MBPer-response size capNo1024
Allow Private IPsPermit the URL to resolve to a private addressNoon
Allow Database OverridePermit a per-query database overrideNoon

Allow Private IPs is on because self-hosted Arc usually runs on a private network or a Docker service name such as http://arc:8000. Turn it off to require a public address. Link-local and cloud-metadata addresses are blocked either way.

Click Save & test to verify the connection.

3. Choosing a protocol

  • Arrow decodes Arc's Arrow IPC stream directly and is the fastest option. Keep it unless you have a reason not to.
  • MessagePack uses Arc's columnar msgpack endpoint. Slightly slower to decode, and it supports gzip compression, which helps on constrained links.
  • JSON is the slowest path, kept for compatibility and debugging.

One practical difference: SHOW DATABASES and SHOW TABLES are not supported on the Arrow endpoint. If you want to use them in a template variable, set the data source to MessagePack or JSON, or query a table directly instead.

4. Get an API token

# Docker — the admin token is printed on first start
docker logs <container-id> 2>&1 | grep "Admin token"

# Or mint a token scoped to Grafana
curl -X POST http://localhost:8000/api/v1/auth/tokens \
  -H "Authorization: Bearer $ARC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "grafana-datasource",
    "description": "Grafana data source access"
  }'

Give Grafana a token scoped to what its dashboard viewers are allowed to read. Arc's token scope, not the plugin, is the authorization boundary.

Enterprise deployments

Which endpoint to point at

A single-node deployment takes the node's address, exactly as in the open-source edition. For a cluster, point the data source at whatever address fronts your query nodes — a load balancer, a Kubernetes service, or a DNS name covering the read replicas. The plugin holds a connection pool per data source and reuses it, so a stable address in front of the cluster gives better connection reuse than rotating DNS.

If Grafana and Arc sit on the same private network, leave Allow Private IPs on. It is on by default for exactly this case.

RBAC: scope the token, not the dashboard

Arc Enterprise enforces role-based access control on every query, per database and per measurement. The plugin sends whatever token the data source is configured with, so the token is what decides which tables a dashboard can read.

That has a practical consequence worth designing around: Grafana's own permissions control who can open a dashboard, not what the SQL inside it may touch. A dashboard editor can write arbitrary SQL, and Arc will run it with the data source's token. So:

  • Create one data source per audience, each with a token scoped to the measurements that audience may read, rather than one shared admin token.
  • Give the Grafana token read-only permissions. Nothing in the plugin writes.
  • Keep Allow Database Override off on a data source whose token spans more databases than its dashboards should reach. With it off, a panel cannot point itself at another database; with it on, it can, within whatever the token permits.

RBAC checks run on every query path, and the check accounts for the database header the plugin sends, so a panel cannot reach another database by changing that header alone.

When RBAC is not licensed, these checks do not apply and a valid token reads everything. Scope tokens accordingly on a deployment that has not enabled it.

Query governance

Where query governance is enabled, Arc may reject or cancel a query that exceeds its limits. Grafana surfaces that as a failed panel with Arc's message. If dashboards fail under load rather than on their SQL, check the governance limits before tuning the plugin — the plugin's own concurrency settings shape how many requests arrive, not what Arc will accept.

Concurrency against a cluster

Max In Flight bounds simultaneous requests from one data source across every panel and viewer; Max Concurrency bounds the chunk fan-out within a single split query. On a cluster serving several Grafana instances, the total load is the sum across data sources, so set these with the cluster's capacity in mind rather than per dashboard.

Writing queries

Referring to tables

The data source's Database setting is sent with every query, so write table names unqualified:

-- Correct: the data source is configured with Database = telegraf
SELECT time, usage_idle FROM cpu WHERE $__timeFilter(time)
-- Rejected: "Cross-database queries (db.table syntax) not allowed
-- when x-arc-database header is set"
SELECT time, usage_idle FROM telegraf.cpu WHERE $__timeFilter(time)

To query a different database from one panel, use the Database field in the query editor rather than a qualified name. That requires Allow Database Override on the data source.

Basic query

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(usage_idle) * -1 + 100 AS cpu_usage,
  host
FROM cpu
WHERE cpu = 'cpu-total'
  AND $__timeFilter(time)
GROUP BY 1, host
ORDER BY time ASC

Set the panel's Format to Time series so Grafana treats the first column as the time axis.

Macros

MacroExpands toExample
$__timeFilter(column)A range predicate on the columnWHERE $__timeFilter(time)
$__timeFrom()Start of the dashboard rangetime >= $__timeFrom()
$__timeTo()End of the dashboard rangetime < $__timeTo()
$__intervalAn interval sized from the rangetime_bucket(INTERVAL '$__interval', time)
$__interval_msThe same interval in millisecondsSELECT $__interval_ms
$__timeGroup(column, interval)A timezone-aware time bucket$__timeGroup(time, '1d')

How they expand:

-- You write
WHERE $__timeFilter(time)

-- Arc receives
WHERE (time) >= '2026-01-17T10:00:00Z' AND (time) < '2026-01-17T11:00:00Z'

$__timeGroup and timezones

Arc stores and returns timestamps in UTC, and Grafana renders them in the dashboard's timezone. Anything grouped by day or larger has to bucket in that timezone too — otherwise a "day" starts at 00:00 UTC, which in UTC−6 is 18:00 the previous evening, and every bar mixes two local calendar days.

$__timeGroup buckets by the dashboard's timezone, so a daily bucket is the viewer's day rather than a UTC day:

SELECT
  $__timeGroup(time, '1d') AS time,
  COUNT(*) AS samples
FROM cpu
WHERE $__timeFilter(time)
GROUP BY 1
ORDER BY time ASC

$__timeGroup accepts any <n><unit> interval — 20s, 2m, 10 minutes, 1h, 1d, 1w — in short or long form. An interval it cannot parse is left unexpanded, so Arc returns a clear error rather than silently bucketing differently.

In a UTC−6 dashboard those buckets start at 06:00 UTC, which is local midnight. Setting the dashboard to "Browser Time" means each viewer sees their own days.

Three details worth knowing:

  • UTC dashboards are unaffected. They use the same epoch arithmetic every earlier release used.
  • Hour buckets stay on epoch arithmetic in any zone whose offset is a whole number of hours, which is nearly all of them. The result is identical, and it avoids a daylight-saving hazard where a repeated local hour would merge two buckets into one.
  • 7d is not a calendar week. Week buckets anchor on Monday, so ask for 1w if that is what you mean. 7d stays a fixed seven-day span.

$__timeGroup also disables query splitting when the bucket is wider than a chunk, because a bucket split across chunks would come back as several partial rows.

More examples

Memory:

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(used_percent) AS memory_used,
  host
FROM mem
WHERE $__timeFilter(time)
GROUP BY 1, host
ORDER BY time ASC

Network throughput, bytes to bits:

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(bytes_recv) * 8 AS bits_in,
  AVG(bytes_sent) * 8 AS bits_out,
  host,
  interface
FROM net
WHERE $__timeFilter(time)
GROUP BY 1, host, interface
ORDER BY time ASC

Disk I/O:

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(read_bytes) AS disk_read,
  AVG(write_bytes) AS disk_write,
  host
FROM diskio
WHERE $__timeFilter(time)
GROUP BY 1, host
ORDER BY time ASC

Template variables

Creating a variable

  1. Dashboard settings → Variables → Add variable
  2. Choose Query, select the Arc data source, and write SQL returning one column
-- Host variable
SELECT DISTINCT host FROM cpu ORDER BY host
-- Interface variable
SELECT DISTINCT interface FROM net ORDER BY interface

Avoid SHOW TABLES in a variable query unless the data source uses the MessagePack or JSON protocol; the Arrow endpoint rejects it.

Quoting rules

This plugin follows the same rule as Grafana's built-in SQL data sources, so queries port between them unchanged:

Variable kindWhat the plugin doesHow to write it
Single-valueEscapes embedded quotes; adds no quotesYou supply them: WHERE host = '$server'
Multi-value or "Include All"Quotes each value and joins with commasLeave it bare: WHERE host IN ($servers)

So a single-value variable works inside a literal of any shape, including a regex:

WHERE host = '$server'
WHERE host ~ '^$server$'

and a multi-value variable must not be wrapped in quotes:

-- Correct
WHERE host IN ($servers)

-- Wrong: produces ''a','b''
WHERE host IN ('$servers')

If you need a multi-value variable inside a regex literal, use ${servers:regex}.

Embedded single quotes are always doubled, so a value cannot terminate the literal it sits in. Two cases fall outside that guarantee, both identical to Grafana's own SQL data sources: a variable used without quotes (WHERE host = $server), and a variable inside a DuckDB escape-string literal (E'$server'). Both are reasons to scope Arc's API token to what dashboard viewers may read.

Using a variable

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(usage_idle) * -1 + 100 AS cpu_usage
FROM cpu
WHERE host = '$server'
  AND cpu = 'cpu-total'
  AND $__timeFilter(time)
GROUP BY 1
ORDER BY time ASC

Alerting

Grafana alert rules work against Arc queries. Because an alert has no dashboard, it has no timezone: $__timeGroup buckets in UTC on the alerting path.

Creating a rule

  1. Open a panel with an Arc query
  2. Alert tab → New alert rule
  3. Set the condition and evaluation interval

Example: high CPU

SELECT
  time,
  100 - usage_idle AS cpu_usage,
  host
FROM cpu
WHERE cpu = 'cpu-total'
  AND time >= NOW() - INTERVAL '5 minutes'
ORDER BY time ASC

Condition: WHEN avg() OF query(A, 5m, now) IS ABOVE 80

Example: memory pressure

SELECT
  time,
  used_percent AS memory_used,
  host
FROM mem
WHERE time >= NOW() - INTERVAL '5 minutes'
ORDER BY time ASC

Condition: WHEN avg() OF query(A, 5m, now) IS ABOVE 90

Route rules to a contact point under Alerting → Contact points.

Dashboard examples

CPU by host (time series):

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  AVG(100 - usage_idle) AS cpu_usage,
  host
FROM cpu
WHERE cpu = 'cpu-total' AND $__timeFilter(time)
GROUP BY 1, host
ORDER BY time ASC

Disk usage (gauge):

SELECT
  host,
  AVG(used_percent) AS disk_used
FROM disk
WHERE $__timeFilter(time)
GROUP BY host

Top hosts by CPU (bar gauge):

SELECT
  host,
  AVG(100 - usage_idle) AS avg_cpu
FROM cpu
WHERE cpu = 'cpu-total' AND $__timeFilter(time)
GROUP BY host
ORDER BY avg_cpu DESC
LIMIT 10

Advanced queries

Moving average (window function):

SELECT
  time,
  usage_idle,
  host,
  AVG(usage_idle) OVER (
    PARTITION BY host
    ORDER BY time
    ROWS BETWEEN 5 PRECEDING AND CURRENT ROW
  ) AS moving_avg
FROM cpu
WHERE cpu = 'cpu-total' AND $__timeFilter(time)
ORDER BY time ASC

Percentiles:

SELECT
  time_bucket(INTERVAL '$__interval', time) AS time,
  host,
  quantile_cont(usage_idle, 0.50) AS p50,
  quantile_cont(usage_idle, 0.95) AS p95,
  quantile_cont(usage_idle, 0.99) AS p99
FROM cpu
WHERE cpu = 'cpu-total' AND $__timeFilter(time)
GROUP BY 1, host
ORDER BY time ASC

Queries containing a window function are not split into chunks, since a window spanning a chunk boundary would produce wrong results.

Performance

Query splitting breaks a long range into chunks that run in parallel. It is skipped automatically where chunking would change results: LIMIT queries, aggregations without a time bucket, UNION, window functions, buckets wider than a chunk, and queries with no time filter to split along.

Keep the range honest. $__timeFilter() is what lets Arc prune files; a panel without it scans everything.

Let $__interval size the buckets. It is derived from the selected range, so the point count stays sensible as the range grows:

-- Good
time_bucket(INTERVAL '$__interval', time)

-- Bad: a day of data at one-second resolution
time_bucket(INTERVAL '1 second', time)

Tune concurrency for the deployment. Max Concurrency bounds one query's chunk fan-out; Max In Flight bounds the whole data source across every panel and viewer. Lower them to protect a busy Arc, raise Max In Flight if a large dashboard's panels queue behind each other.

Cache repeated queries. Grafana's per-data-source query caching helps most on dashboards several people watch at once.

Troubleshooting

The data source does not appear

ls -la /var/lib/grafana/plugins/basekick-arc-datasource
grep -i "unsigned\|basekick" /var/log/grafana/grafana.log

The usual cause is the unsigned-plugin setting. Grafana logs Plugin is unsigned and skips loading unless allow_loading_unsigned_plugins names basekick-arc-datasource.

Save & test fails

# Is Arc up?
curl http://localhost:8000/health

# Is the token valid?
curl -H "Authorization: Bearer $ARC_TOKEN" http://localhost:8000/api/v1/auth/verify

If the message mentions a blocked address, the URL resolves to a private range and Allow Private IPs is off.

A query fails

Parser and syntax errors from DuckDB are shown on the panel. Other errors are summarised, with the full text and the expanded SQL in the Grafana server log:

grep "Arc query failed" /var/log/grafana/grafana.log | tail

Common causes:

  • "Cross-database queries not allowed" — a db.table name while the data source has a Database configured. Drop the prefix.
  • "SHOW DATABASES is not supported on the Arrow endpoint" — switch the data source to MessagePack or JSON, or query a table instead.
  • "Table with name X does not exist" — the table is in a different database than the one configured.

Listing what exists

SHOW statements need the MessagePack or JSON protocol, and Arc runs one statement per query, so use them in separate panels or variable queries:

SHOW DATABASES
SHOW TABLES

Slow dashboards

Most panel latency is Arc-side query time, not plugin overhead. Check a query directly:

curl -X POST http://localhost:8000/api/v1/query \
  -H "Authorization: Bearer $ARC_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"sql": "SELECT count(*) FROM cpu WHERE time > NOW() - INTERVAL 1 HOUR"}'

If a single query is fast but the dashboard is slow, raise Max In Flight so panels stop queuing.

Resources

Next steps

On this page