Publication timing can be hard to see
Publication schedules and update timing may be documented in different places. DataPulse records observable timing signals alongside the published dataset.
The trust layer for Malaysian open data
Malaysia ranks #1 in the world for open data (ODIN 2024/25, 99/100 for openness) — and freshness is the gap ODIN doesn't measure. We monitor official datasets. The scheduler wakes every 5 minutes and probes only datasets due under their tiered cadence, classify each into one of ten honest statuses, and expose it read-only over MCP.
Verified by mcpgrade: 100/100 (Grade A), 16 tools, last audited 2026-08-17.
No API key. Read-only. Works with Claude Desktop, Cursor, Cline, and any MCP client.
{
"mcpServers": {
"datapulse-my": {
"transport": "streamable-http",
"url": "https://mcp.data-pulse.my/mcp"
}
}
}
search_datasetsget_datasetfind_stalefind_anomaliesfind_deterioratingfind_recoveringfind_unreliablefind_schema_driftcheck_reconciliationget_provenancefind_by_licence$ curl -s https://mcp.data-pulse.my/mcp ...
# → Read-only catalogue context with provenance and publication signals.
Public-service companion layer
Malaysian agencies and public institutions publish data for public use. DataPulse adds independent, read-only observations of availability, freshness signals, provenance, and licence or reuse context; official publishers remain the source of record.
Publication schedules and update timing may be documented in different places. DataPulse records observable timing signals alongside the published dataset.
Sources expose different combinations of headers, content dates, and page information. DataPulse records which observable signals are available for review.
Licence and reuse details can accompany a catalogue entry rather than the downloaded file. DataPulse preserves the available context in machine-readable form.
How DataPulse supports review
Each cycle documents observable signals through the same inspectable path, supporting review alongside the publisher’s own information.
A 5-minute scheduler wakes and probes datasets due under their tiered cadence: HTTP headers, content dates, record counts, schema shape.
Three freshness signals: Last-Modified header, content-date parse, browser extraction when needed.
One of ten statuses based on the available observations. When freshness cannot be established from those signals, that evidence gap is recorded.
Health JSON, trend reliability evidence, badge, report, RSS, and MCP tools update atomically — all machine-readable.
A complementary layer around official data
Official publishers remain the source of record. DataPulse contributes read-only, independently observed context that can support review, documentation, and reuse.
| Context | What DataPulse adds |
|---|---|
| Publication signals | Observed timestamps and availability signals, with explicit evidence gaps where signals are unavailable. |
| Provenance and reuse | Machine-readable provenance and licence or reuse context that accompanies the catalogue record. |
| Agent access | Read-only catalogue context for agents, without changing the published source. |
| Structural observations | Published observations of schema shape, record counts, history, and dataset deltas. |
DataPulse documents observed context and evidence gaps; it does not replace the published source.
5 of 389 datasets (1.3%) require a real browser to probe because their
source pages render client-side JavaScript: eperolehan-diklankan,
doe_apims, doe_rqims, doe_mqims, and kkm_idengue.
DataPulse uses Camofox,
a self-hosted patched headless-Chromium sidecar, to probe these. The probe path
is check.sh → Camofox sidecar → DOM snapshot → content-date extraction.
To enable browser probing: set CAMOFOX_BASE_URL and restart the timer.
Without Camofox, those 5 datasets sit at browser-dependent — the honest status
when no browser is available.
DataPulse probes publicly-published open-data sources. We do not bypass authentication, CAPTCHAs, or terms-of-service restrictions. Every source we probe is publicly available without login; the data is aggregate/non-personal; and the probe respects each dataset's declared refresh frequency.
All scraping is rate-limited (5-minute cadence, dataset-tier cadence applied)
and identifies itself via User-Agent. Sources we cannot probe without authentication,
CAPTCHA bypass, or ToS violation are marked unreachable or
browser-dependent — never silently scraped through a workaround.
We also respect each source's robots.txt via respect_robots_txt()
in scripts/check.sh — sources that explicitly disallow automated
access are not probed.
If you are a data source maintainer and would like DataPulse to adjust its probe cadence, exclude a dataset, or remove it from the manifest, please open a GitHub issue.
Loading signal gaps…
FAQ
Yes — Datasets are probed when due under their tiered cadence. See the live status above.
No — it verifies and documents official sources. We never resell the data.
Within 1.5× the dataset’s declared cadence, verified by Last-Modified header or content-date.
Yes — the MCP server is read-only and key-free.
Yes — MIT, on GitHub, with public dataset deltas and reproducible build proof.
Give your agent verified Malaysian public-data context in one connection.
Live from the latest committed health envelope
Loading live dataset health…