Skip to main content
Loading news...

Space Frontiers research corpus

Space Frontiers is built on a single continuously updated corpus: scholarly literature, books, patents, encyclopedias, and public community discussions, collected, parsed, and embedded for retrieval. This page describes what is inside.

Production database snapshot · July 22, 2026

documents
2.82B
tokens, estimated
≈1.8T
scholarly & reference works
195M
languages
50+
publication span
1840s–2026
text & vectors on disk
24.5 TB

Access, provenance & rights

Search access and bulk data access are separate products. You can inspect records through the web app, REST API, or MCP server before discussing a licensed bulk delivery.

Search access
Results include canonical source URIs and available bibliographic metadata. Full text is returned only where it is present in the corpus.
Bulk access
Scoped by source segment, intended use, and volume. There is no anonymous whole-corpus download.
Provenance
DOI, PubMed, ISBN, patent, and source-page identifiers are retained where available so records can be traced back to their source.
Rights
Availability and permitted use vary by source and content segment. A Space Frontiers delivery does not replace third-party rights or attribution requirements.

Scholarly & reference works

The core of the corpus: 195M scholarly and reference records with metadata, abstracts, and full text where available.

TypeRecordsShare
Journal articles132M67%
Book chapters23.2M12%
Proceedings articles9.4M4.8%
Encyclopedia articles9.3M4.8%
Books & monographs6.1M3.1%
Patents5.1M2.6%
Components (figures, data, supplements)4.7M2.4%
Preprints & posted content3.3M1.7%
Reports700K0.4%
Reference entries350K0.2%
Standards, dissertations & grants140K0.1%
Other1.5M0.8%

Full text and abstracts

  • Journal articles with abstract79% · 104M
  • Patents with full text100% · 5.1M
  • Encyclopedia articles with full text100% · 9.3M
  • Books with complete text61% · 3.7M
  • Journal articles with full text32% · 42M

Linked identifiers

DOI-linked works
158M
PubMed / PMC
46M
ISBN
11M
arXiv
1.4M

A 0.5% random sample of articles alone spans 16,000+ distinct journals and 3,300+ publishers, so full journal coverage is substantially wider. Publication dates reach back to the 1840s.

Community discussions

Public conversation streams complement the scholarly record with current, informal knowledge.

SourceDocuments Communities
Telegram public channels 176.7M messages 194,905 channels
Reddit submissions 2.44B posts 14,000+ subreddits
Reddit comments 23.5B comments

Text volume

Estimated token counts by segment, assuming ≈4 characters per token. The corpus totals roughly 7.3 trillion characters, or ≈1.8 trillion tokens.

SegmentEst. tokens
Books & monographs≈345B
Journal articles & abstracts≈280B
Patents≈58B
Book chapters & proceedings≈25B
Encyclopedia articles≈24B
Other scholarly text≈5B
Telegram messages≈21B
Reddit posts & comments≈1.06T

Languages

50+ languages appear in language-tagged content (scholarly works and Telegram messages). Untagged Reddit content is predominantly English.

  • English56%
  • Russian30%
  • Arabic1.9%
  • German1.7%
  • Persian1.5%
  • Spanish1.5%
  • French1.1%
  • Ukrainian0.7%
  • Portuguese0.5%
  • Other (40+)5.1%

Publication dates

Scholarly coverage is deep — a third of works were published in the 2020s, yet the corpus reaches back past 1900. Community content concentrates in recent years: 63% was published within the last five years.

Scholarly works, by decade

  • pre-19004.8%
  • 1900–19494.2%
  • 1950s2.2%
  • 1960s1.7%
  • 1970s3%
  • 1980s5.5%
  • 1990s6.8%
  • 2000s12.6%
  • 2010s26.3%
  • 2020s33.1%

Community content, by year

  • pre-20143.3%
  • 20141.8%
  • 20152.3%
  • 20162.8%
  • 20173.1%
  • 20184.3%
  • 20196.4%
  • 20208%
  • 20218.1%
  • 20229.4%
  • 202314.6%
  • 202414.7%
  • 202514.6%
  • 2026*6.4%

* through July 2026

Embeddings & storage

Every document is chunked and embedded for hybrid retrieval: 3.48 billion passages carry both dense 2,560-dimension binary-quantized vectors and learned sparse vectors over a ≈106,000-term vocabulary. 82% of passages embed full document content, 13% condensed short-document views, and 5% abstracts.

On-disk footprint Size (LZ4-compressed)
Documents & metadata8.7 TB
Reddit comments7.9 TB
Embeddings5.1 TB
Identifiers, queues & auxiliary2.8 TB
Total24.5 TB

Methodology

Figures come from block-level random samples and planner statistics taken against the production database on July 22, 2026. Token counts are estimated at ≈4 characters per token; journal and publisher counts are lower bounds observed in samples. All values are rounded and refreshed periodically as the corpus grows.

Bulk access & licensing

The corpus is available for licensing — for research, AI training, and enterprise applications. Tell us which segments you need, your intended use case, and expected volume, and we will suggest a delivery format.

Contact us about dataset access