Supported Data Sources

The Dubstrata Engine ingests, processes, and structures alternative data from a wide array of public, prediction, and financial sources to build a point-in-time causal knowledge graph.

Prediction Markets & Orderbooks

Direct integrations with decentralized prediction markets to capture crowdsourced probabilities and sentiment spikes.

  • Polymarket Gamma API (Market metadata, outcomes, contract IDs, and resolution updates)
  • Polymarket Central Limit Order Book (CLOB) (Real-time order books, transaction volume, depth charts, and whale concentration)

Financial Price Feeds

Global spot pricing feeds mapped to digital assets and macro indices to compute correlation metrics.

  • Crypto asset tick feeds (24h trade metrics and spot pricing anomalies)
  • Global macro price feeds (Equities, indices, commodities, and currency exchange rates)

Reference Registries & Search

Information sources leveraged by the self-reflective JIT engine to populate background facts and entity definitions.

  • Wikipedia API (Used for foundational definitions, historical events, and corporate profiles)
  • DuckDuckGo Search Integration (Utilized by background crawlers to locate real-time updates and news articles on demand)

Public Publications & Feeds

Official gazettes and news feeds scraped to verify regulatory statements and track trends.

  • Sovereign Regulatory Gazettes (Official filings, regulatory announcements, and enforcement releases)
  • Passive RSS News Feeds (Financial, macroeconomic, and sector-specific news feeds updated on a 4-hour schedule)

Multi-Tenant Scope Boundaries

While these public feeds populate the default global schema, all tenant-specific documents ingested via the ingest_knowledge tool (with is_private: true) are isolated into dedicated, encrypted namespaces. These private sources are completely excluded from other tenants' queries.