Space Frontiers research corpus
Space Frontiers is built on a single continuously updated corpus: scholarly literature, books, patents, encyclopedias, and public community discussions, collected, parsed, and embedded for retrieval. This page describes what is inside.
Production database snapshot · July 22, 2026
- documents
- 2.82B
- tokens, estimated
- ≈1.8T
- scholarly & reference works
- 195M
- languages
- 50+
- publication span
- 1840s–2026
- text & vectors on disk
- 24.5 TB
Access, provenance & rights
Search access and bulk data access are separate products. You can inspect records through the web app, REST API, or MCP server before discussing a licensed bulk delivery.
- Search access
- Results include canonical source URIs and available bibliographic metadata. Full text is returned only where it is present in the corpus.
- Bulk access
- Scoped by source segment, intended use, and volume. There is no anonymous whole-corpus download.
- Provenance
- DOI, PubMed, ISBN, patent, and source-page identifiers are retained where available so records can be traced back to their source.
- Rights
- Availability and permitted use vary by source and content segment. A Space Frontiers delivery does not replace third-party rights or attribution requirements.
Scholarly & reference works
The core of the corpus: 195M scholarly and reference records with metadata, abstracts, and full text where available.
| Type | Records | Share of scholarly records | Share |
|---|---|---|---|
| Journal articles | 132M | 67% | |
| Book chapters | 23.2M | 12% | |
| Proceedings articles | 9.4M | 4.8% | |
| Encyclopedia articles | 9.3M | 4.8% | |
| Books & monographs | 6.1M | 3.1% | |
| Patents | 5.1M | 2.6% | |
| Components (figures, data, supplements) | 4.7M | 2.4% | |
| Preprints & posted content | 3.3M | 1.7% | |
| Reports | 700K | 0.4% | |
| Reference entries | 350K | 0.2% | |
| Standards, dissertations & grants | 140K | 0.1% | |
| Other | 1.5M | 0.8% |
Full text and abstracts
- Journal articles with abstract79% · 104M
- Patents with full text100% · 5.1M
- Encyclopedia articles with full text100% · 9.3M
- Books with complete text61% · 3.7M
- Journal articles with full text32% · 42M
Linked identifiers
- DOI-linked works
- 158M
- PubMed / PMC
- 46M
- ISBN
- 11M
- arXiv
- 1.4M
A 0.5% random sample of articles alone spans 16,000+ distinct journals and 3,300+ publishers, so full journal coverage is substantially wider. Publication dates reach back to the 1840s.
Community discussions
Public conversation streams complement the scholarly record with current, informal knowledge.
| Source | Documents | Communities |
|---|---|---|
| Telegram public channels | 176.7M messages | 194,905 channels |
| Reddit submissions | 2.44B posts | 14,000+ subreddits |
| Reddit comments | 23.5B comments |
Text volume
Estimated token counts by segment, assuming ≈4 characters per token. The corpus totals roughly 7.3 trillion characters, or ≈1.8 trillion tokens.
| Segment | Est. tokens | Share of tokens |
|---|---|---|
| Books & monographs | ≈345B | |
| Journal articles & abstracts | ≈280B | |
| Patents | ≈58B | |
| Book chapters & proceedings | ≈25B | |
| Encyclopedia articles | ≈24B | |
| Other scholarly text | ≈5B | |
| Telegram messages | ≈21B | |
| Reddit posts & comments | ≈1.06T |
Languages
50+ languages appear in language-tagged content (scholarly works and Telegram messages). Untagged Reddit content is predominantly English.
- English56%
- Russian30%
- Arabic1.9%
- German1.7%
- Persian1.5%
- Spanish1.5%
- French1.1%
- Ukrainian0.7%
- Portuguese0.5%
- Other (40+)5.1%
Publication dates
Scholarly coverage is deep — a third of works were published in the 2020s, yet the corpus reaches back past 1900. Community content concentrates in recent years: 63% was published within the last five years.
Scholarly works, by decade
- pre-19004.8%
- 1900–19494.2%
- 1950s2.2%
- 1960s1.7%
- 1970s3%
- 1980s5.5%
- 1990s6.8%
- 2000s12.6%
- 2010s26.3%
- 2020s33.1%
Community content, by year
- pre-20143.3%
- 20141.8%
- 20152.3%
- 20162.8%
- 20173.1%
- 20184.3%
- 20196.4%
- 20208%
- 20218.1%
- 20229.4%
- 202314.6%
- 202414.7%
- 202514.6%
- 2026*6.4%
* through July 2026
Embeddings & storage
Every document is chunked and embedded for hybrid retrieval: 3.48 billion passages carry both dense 2,560-dimension binary-quantized vectors and learned sparse vectors over a ≈106,000-term vocabulary. 82% of passages embed full document content, 13% condensed short-document views, and 5% abstracts.
| On-disk footprint | Size (LZ4-compressed) |
|---|---|
| Documents & metadata | 8.7 TB |
| Reddit comments | 7.9 TB |
| Embeddings | 5.1 TB |
| Identifiers, queues & auxiliary | 2.8 TB |
| Total | 24.5 TB |
Methodology
Figures come from block-level random samples and planner statistics taken against the production database on July 22, 2026. Token counts are estimated at ≈4 characters per token; journal and publisher counts are lower bounds observed in samples. All values are rounded and refreshed periodically as the corpus grows.
Bulk access & licensing
The corpus is available for licensing — for research, AI training, and enterprise applications. Tell us which segments you need, your intended use case, and expected volume, and we will suggest a delivery format.
Contact us about dataset access