Space Frontiers research corpus
Space Frontiers is built on a single continuously updated corpus: scholarly literature, books, patents, encyclopedias, and public community discussions, collected, parsed, and embedded for retrieval. This page describes what is inside.
Production database snapshots · July 22–August 13, 2026
- documents
- 2.82B
- tokens, estimated
- ≈1.8T
- scholarly & reference works
- 195M
- original source files
- 76.9M
- source file archive
- 128.9 TB
- languages
- 100+
- publication span
- 1840s–2026
- text & vectors on disk
- 24.5 TB
Access, provenance & rights
Search access and bulk data access are separate products. You can inspect records through the web app, REST API, or MCP server before discussing a licensed bulk delivery.
- Search access
- Results include canonical source URIs and available bibliographic metadata. Full text is returned only where it is present in the corpus.
- Bulk access
- Scoped by source segment, intended use, and volume. There is no anonymous whole-corpus download.
- Provenance
- DOI, PubMed, ISBN, patent, and source-page identifiers are retained where available so records can be traced back to their source.
- Rights
- Availability and permitted use vary by source and content segment. A Space Frontiers delivery does not replace third-party rights or attribution requirements.
Scholarly & reference works
The core of the corpus: 195M scholarly and reference records with metadata, abstracts, and full text where available.
| Type | Records | Share of scholarly records | Share |
|---|---|---|---|
| Journal articles | 132M | 67% | |
| Book chapters | 23.2M | 12% | |
| Proceedings articles | 9.4M | 4.8% | |
| Encyclopedia articles | 9.3M | 4.8% | |
| Books & monographs | 6.1M | 3.1% | |
| Patents | 5.1M | 2.6% | |
| Components (figures, data, supplements) | 4.7M | 2.4% | |
| Preprints & posted content | 3.3M | 1.7% | |
| Reports | 700K | 0.4% | |
| Reference entries | 350K | 0.2% | |
| Standards, dissertations & grants | 140K | 0.1% | |
| Other | 1.5M | 0.8% |
Full text and abstracts
- Journal articles with abstract79% · 104M
- Patents with full text100% · 5.1M
- Encyclopedia articles with full text100% · 9.3M
- Books with complete text61% · 3.7M
- Journal articles with full text32% · 42M
Linked identifiers
- DOI-linked works
- 158M
- PubMed / PMC
- 46M
- ISBN
- 11M
- arXiv
- 1.4M
A 0.5% random sample of articles alone spans 16,000+ distinct journals and 3,300+ publishers, so full journal coverage is substantially wider. Publication dates reach back to the 1840s.
Community discussions
Public conversation streams complement the scholarly record with current, informal knowledge.
| Source | Documents | Communities |
|---|---|---|
| Telegram public channels | 176.7M messages | 194,905 channels |
| Reddit submissions | 2.44B posts | 14,000+ subreddits |
| Reddit comments | 23.5B comments |
Text volume
Estimated token counts by segment, assuming ≈4 characters per token. The corpus totals roughly 7.3 trillion characters, or ≈1.8 trillion tokens.
| Segment | Est. tokens | Share of tokens |
|---|---|---|
| Books & monographs | ≈345B | |
| Journal articles & abstracts | ≈280B | |
| Patents | ≈58B | |
| Book chapters & proceedings | ≈25B | |
| Encyclopedia articles | ≈24B | |
| Other scholarly text | ≈5B | |
| Telegram messages | ≈21B | |
| Reddit posts & comments | ≈1.06T |
Languages
100+ languages appear in language-tagged documents. This distribution uses only scholarly and reference records; Telegram, Reddit, Discord, and YouTube records are excluded.
- English84.2%
- German3.8%
- French2.2%
- Russian1.8%
- Spanish1.4%
- Portuguese1.1%
- Italian0.5%
- Chinese0.5%
- Japanese0.4%
- Other (90+)4.1%
Publication dates
Scholarly coverage is deep — a third of works were published in the 2020s, yet the corpus reaches back past 1900. Community content concentrates in recent years: 63% was published within the last five years.
Scholarly works, by decade
- pre-19004.8%
- 1900–19494.2%
- 1950s2.2%
- 1960s1.7%
- 1970s3%
- 1980s5.5%
- 1990s6.8%
- 2000s12.6%
- 2010s26.3%
- 2020s33.1%
Community content, by year
- pre-20143.3%
- 20141.8%
- 20152.3%
- 20162.8%
- 20173.1%
- 20184.3%
- 20196.4%
- 20208%
- 20218.1%
- 20229.4%
- 202314.6%
- 202414.7%
- 202514.6%
- 2026*6.4%
* through July 2026
Embeddings & storage
Every document is chunked and embedded for hybrid retrieval: 3.48 billion passages carry both dense 2,560-dimension binary-quantized vectors and learned sparse vectors over a ≈106,000-term vocabulary. 82% of passages embed full document content, 13% condensed short-document views, and 5% abstracts.
| On-disk footprint | Size (LZ4-compressed) |
|---|---|
| Documents & metadata | 8.7 TB |
| Reddit comments | 7.9 TB |
| Embeddings | 5.1 TB |
| Identifiers, queues & auxiliary | 2.8 TB |
| Total | 24.5 TB |
Original source files
Separately from the search database, every acquired source document is retained in its original form in object storage: 76.9 million files totalling 128.9 TB. PDFs dominate with 71.9 million files and 119.7 TB, alongside 2.2 million EPUBs, 171,000 DJVU scans, and publisher HTML and JATS XML sources.
Methodology
Most figures come from block-level random samples and planner statistics taken against the production database on July 22, 2026. The language distribution is the rounded mean of two reproducible 0.05% block-level samples of non-social document records taken on August 13, 2026. Original source file counts and volume were measured directly against object storage on August 8, 2026. Token counts are estimated at ≈4 characters per token; journal and publisher counts are lower bounds observed in samples. All values are rounded and refreshed periodically as the corpus grows.
Bulk access & licensing
The corpus is available for licensing — for research, AI training, and enterprise applications. Tell us which segments you need, your intended use case, and expected volume, and we will suggest a delivery format.
Contact us about dataset access