Skip to main content

Crate metadata_gen

Crate metadata_gen 

Source
Expand description

metadata-gen logo

metadata-gen

Front matter in, metadata and SEO meta tags out. YAML, TOML and JSON, with zero unsafe code.

Build Registry Docs OpenSSF Scorecard OpenSSF Best Practices License: Apache-2.0 OR MIT MSRV 1.88.0


§Contents

Getting started

The metadata-gen ecosystem

Library reference

Operational


§Install

§As a Rust library

[dependencies]
metadata-gen = "0.0.8"

Or from the command line:

cargo add metadata-gen

There is no CLI binary. metadata-gen is a library crate and ships no [[bin]].

§Build from source

git clone https://github.com/sebastienrousseau/metadata-gen.git
cd metadata-gen
make          # check + clippy + test

§Cargo features

Every format and integration is a feature, and all of them are on by default, so a default build offers everything 0.0.7 did. Turn off what you do not parse with default-features = false.

FeatureDefaultEnables
stdyesFile helpers in metadata_gen::io, std::io::Error in MetadataError, HashMap-backed MetadataMap
yamlyes--- front matter (noyalib)
tomlyes+++ front matter (toml)
jsonyesJSON-object front matter (serde_json)
htmlyesextract_meta_tags (quick-xml); implies std
tokioyesmetadata_gen::tokio and async_extract_metadata_from_file; implies std
# YAML only, no standard library:
metadata-gen = { version = "0.0.8", default-features = false, features = ["yaml"] }

A fence whose format is not compiled in is reported as UnsupportedFormatError, not as missing front matter.


§Requirements

  • Rust 1.88.0 or newer. rust-version in Cargo.toml is the floor and Cargo enforces it; CI builds on stable across Linux, macOS and Windows.
  • alloc, with std optional. With default-features = false the crate is no_std + alloc and MetadataMap is a BTreeMap; cargo build --no-default-features --features yaml --target thumbv7em-none-eabihf builds.
  • No async runtime is required. Every synchronous entry point works without one. The tokio feature adds metadata_gen::tokio for callers who run Tokio; the dependency is trimmed to fs and io-util.

§Quick Start

use metadata_gen::extract_and_prepare_metadata;

let content = "---\n\
title: Hello, world!\n\
description: A short greeting\n\
keywords: rust, frontmatter, seo\n\
---\n\

let (metadata, keywords, tags) =
    extract_and_prepare_metadata(content).expect("valid front matter");

assert_eq!(metadata.get("title"), Some(&"Hello, world!".to_string()));
assert_eq!(keywords, vec!["rust", "frontmatter", "seo"]);
assert!(tags.primary.contains("description"));

extract_and_prepare_metadata detects the format, flattens the metadata into a map, extracts keywords, and generates formatted meta tags in one call.


§The metadata-gen ecosystem

The family releases along the 0.0.x line, with each repository owning a focused role in static site and content pipelines.

ComponentPurposeUse case
metadata-genExtract and process front matter in YAML, TOML, JSON; generate meta tagsContent pipelines, documentation, static site generators
frontmatter-genFront-matter validation and schema generationBuild-time front-matter validation
mdx-genMDX compiler and component rendererEmbedded interactive components in Markdown
sitemap-genXML sitemap generator conforming to sitemaps.orgSearch engine index generation
rssgenRSS, Atom, and JSON feed generatorContent syndication feeds
staticdatagenStatic data engine for template processingTemplate data hydration

§Capabilities at a glance

AreaCapabilityStatus
Front-matter extractionYAML, TOML, and JSON delimited headersStable
Document separationSplit front-matter block and document body in one callStable
Typed extractionDeserialise front matter into user-defined structsStable
Flat metadataFlatten nested structures to dot-separated string mapsStable
Meta tag generationOpen Graph, Twitter, Apple, Microsoft, and primary tagsStable
Meta tag extractionSingle streaming pass over HTML documents with quick-xmlStable
UtilitiesSingle-pass HTML entity escaping and async file loadingStable

§Ecosystem comparison

This summary identifies API shape, not a universal winner. Workload-specific trade-offs and the evidence behind each cell are documented separately.

ProjectYAMLTOMLJSONTypedBody returnedMeta tags
metadata-genYesYesYesYesYesYes
gray_matterYesYesYesYesYesNo
yaml-front-matterYesNoNoYesYesNo
matterYesNoNoNoYesNo

The matrix reflects each crate’s documented surface at the time of this release.


§Benchmarks

Measured with Criterion on an Apple A18 Pro, rustc 1.98.0, single thread. Middle estimate of the confidence interval; run cargo bench to reproduce on your own hardware.

ScenarioResultEnvironment
extract_metadata (YAML, 1 KB)631 µs (1.5 MiB/s)Apple A18 Pro, rustc 1.98.0
extract_metadata (YAML, 10 KB)1.81 ms (5.4 MiB/s)Apple A18 Pro, rustc 1.98.0
extract_metadata (YAML, 1 MB)181 ms (5.5 MiB/s)Apple A18 Pro, rustc 1.98.0
extract_meta_tags (1 KB)50 µs (18.7 MiB/s)Apple A18 Pro, rustc 1.98.0
extract_meta_tags (1 MB)34.9 ms (28.7 MiB/s)Apple A18 Pro, rustc 1.98.0
escape_html (10 KB)31.9 µs (295 MiB/s)Apple A18 Pro, rustc 1.98.0
extract_and_prepare_metadata (~250 B)23 µsApple A18 Pro, rustc 1.98.0

Reproduce with cargo bench --all-features; the harnesses live in benches/.


§Features

  • Three front-matter shapes: YAML (--- … ---), TOML (+++ … +++), and JSON ({ ... } at top of file).
  • Dual extraction APIs:
    • extract_metadata: returns flat Metadata (HashMap<String, String>) with dot-separated keys (author.name), ideal for templates.
    • extract_typed::<T>: hands the raw block to serde, returning your strongly typed struct; extract_typed_borrowed lets &str fields point into the document.
  • Document body access: extract_metadata_with_body returns (Metadata, &str) with the body following the closing delimiter. detect_front_matter exposes format, raw block, and byte offset without parsing.
  • Bounded parsing: every parse runs under ParseLimits (block size, nesting depth); the YAML parser gets its strict budget preset on the flat and typed paths alike. extract_metadata_with_limits takes your own limits.
  • Errors that locate the fault: MetadataError::Parse carries the format, the byte span inside the block and the parser’s error; an unclosed fence reports its byte offset.
  • Processing and normalization: process_metadata and process_metadata_with normalize dates to YYYY-MM-DD (DD/MM/YYYY by default, MM/DD/YYYY with DateOrder::MonthFirst), verify required fields, and derive URL slugs from titles.
  • Meta tag synthesis: generate_metatags creates grouped tags (primary, og, twitter, apple, ms); Open Graph tags use property=, and attribute values pass through escape_attribute. MetaTag and MetaTagGroups::iter give the same tags as values.
  • Streaming HTML tag extraction: extract_meta_tags extracts <meta> tags in a single streaming pass using quick-xml without full DOM allocation; extract_meta_tags_lenient also reports where a malformed page stopped the scan.
  • HTML escaping and file utilities: single-pass, single-allocation escape_html and unescape_html, readers and files in metadata_gen::io, and Tokio-backed helpers in metadata_gen::tokio.

§Configuration

process_metadata uses a default policy: title and date are required, and slug is derived from title. process_metadata_with accepts caller-configured ProcessOptions:

use metadata_gen::{process_metadata_with, Metadata, ProcessOptions};
use std::collections::HashMap;

let options = ProcessOptions::default()
    .required_fields(["title", "author"])
    .derive_slug(false);

let mut map = HashMap::new();
map.insert("title".to_string(), "Hello".to_string());
map.insert("author".to_string(), "Ada".to_string());

let processed = process_metadata_with(&Metadata::new(map), &options).unwrap();
assert!(!processed.contains_key("slug"));
OptionDefaultEffect
required_fields["title", "date"]Missing field returns MissingFieldError naming it
derive_slugtrueDerives slug from title when absent
date_orderDateOrder::DayFirstHow a slash-separated date such as 01/02/2024 is read

ProcessOptions is #[non_exhaustive], so options can be added without breaking releases.


§Examples

Run any example with cargo run --example <name>:

ExampleShows
lib_exampleThe high-level extract_and_prepare_metadata flow
metadata_examplePer-format extraction, nested tables, typed extraction, the body
metatags_exampleGenerating <meta> groups and reading them back
utils_exampleHTML escape/unescape and the async file helper
error_exampleEvery MetadataError variant and how to recover

make examples runs all examples; CI executes the same suite on every push.


§When not to use metadata-gen

  • You need element-level access to arrays of objects from the flat map. [a, b] is rendered as a string in the flat map by design. Use extract_typed::<T> instead, which preserves structure.
  • You need to round-trip front matter byte-for-byte. The crate parses; it does not preserve comments, key order or quoting style, and there is no serialiser back to a fenced block.
  • You need every <meta> element from arbitrary broken HTML. Extraction stops at the first unrecoverable reader error and returns what was found so far. A full HTML5 parser (html5ever, scraper) is the right tool if you need error recovery over broken whole pages.

§Development

make              # check + clippy + test
make test         # all tests, all features
make clippy       # lints, warnings denied
make fmt          # formatting check
make lint         # markdownlint + codespell + REUSE
make doc          # rustdoc with warnings denied
make coverage     # line coverage gate (98%)
make miri         # lib tests under Miri
make proptest     # property-based tests (involution + round-trip)
make loom         # concurrency testing via Loom model checker
make kani         # formal verification proofs via Kani
make mutants      # mutation testing kill rate (cargo-mutants)
make fuzz         # build every target, replay corpus and regressions
make examples     # run every example
make bench-smoke  # compile and run each bench once
make versions     # every version-bearing file agrees
make complexity   # per-function complexity ceilings
make links        # every Markdown link resolves
make msrv         # builds on the declared minimum Rust
make semver       # public API against the last release
make hack         # every feature combination compiles
make distcheck    # package and verify the archive
make deny / vet / audit   # supply chain

DEVELOPMENT.md maps each CI job to its local equivalent and explains reproduction steps. docs/MIGRATION.md lists what changes for consumers between releases.


§Security

Reporting: never open a public issue for a vulnerability. See SECURITY.md for disclosure instructions.

  • #![forbid(unsafe_code)] proves the absence of unsafe code (ADR-0001).
  • No C dependencies, no FFI, no network I/O, no environment reads.
  • Meta-tag attribute values are escaped on generation, preventing markup injection.
  • Supply chain audited via cargo-audit, cargo-deny, and cargo-vet with a locked exemption baseline.
  • Front matter is parsed under ParseLimits (block size and depth), with the YAML parser’s strict resource budgets on every path, the typed one included.
  • The first-party dependency noyalib is pinned exactly (ADR-0004).
  • Property-based testing via Proptest for parser round-trips and HTML escape involution.
  • Concurrency testing scaffold via Loom (tests/loom_smoke.rs) for thread schedule exploration.
  • Formal verification via Kani (tests/kani/) proving HTML escape totality and ASCII round-trip.
  • Mutation testing via cargo-mutants configured with .cargo/mutants.toml.
  • Three fuzz targets replaying seed and regression corpora per push.

§Documentation

The canonical entry points across the repository family:

DocumentCovers
CHANGELOG.mdPer-release notes, Keep a Changelog format
SECURITY.mdDisclosure policy, supported versions, security design
CONTRIBUTING.mdBranch and commit conventions, PR expectations
GOVERNANCE.mdProject stewardship, how changes land
SUPPORT.mdSupport channels and expectations
AGENTS.mdInvariants for AI-assisted contributions

§Stability guarantees

§Versioning

SemVer 2.0.0 is followed, with the pre-1.0 posture that releases increment strictly by +0.0.1 along the 0.0.x line. Every breaking change is documented in CHANGELOG.md.

§Output stability

What the crate produces is part of the API: a change to how a document flattens, which key a value lands under, or what a meta-tag group renders is treated as a breaking change even when no Rust signature moves.

§Deprecations

Deprecations live for at least two releases with a #[deprecated] attribute naming the replacement before removal.

§Minimum-toolchain discipline

The floor is Rust 1.88.0, declared as rust-version in Cargo.toml. Raising it is a breaking change and occurs only on releases with rationale documented in the changelog. See docs/POLICIES.md for the complete policy.

§Version-bearing files

Checked against the manifest by scripts/verify-release-versions.sh before a tag exists, ensuring install snippets and metadata stay synchronised.


§License

Dual-licensed under Apache 2.0 or MIT, at your option.

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this crate by you shall be dual-licensed as above, without any additional terms or conditions.

Back to top

Re-exports§

pub use error::MetadataError;
pub use metadata::detect_front_matter;
pub use metadata::extract_metadata;
pub use metadata::extract_metadata_with_body;
pub use metadata::extract_metadata_with_limits;
pub use metadata::extract_typed;
pub use metadata::extract_typed_borrowed;
pub use metadata::process_metadata;
pub use metadata::process_metadata_with;
pub use metadata::DateOrder;
pub use metadata::FrontMatterFormat;
pub use metadata::Metadata;
pub use metadata::ParseLimits;
pub use metadata::ProcessOptions;
pub use metatags::generate_metatags;
pub use metatags::MetaTag;
pub use metatags::MetaTagGroups;
pub use utils::async_extract_metadata_from_file;tokio and non-loom
pub use utils::escape_attribute;
pub use utils::escape_html;

Modules§

error
The error module contains error types for metadata processing. Error types for the metadata-gen library.
iostd
Synchronous readers and files (std). Synchronous readers and files.
metadata
The metadata module contains functions for extracting and processing metadata. Metadata extraction and processing module.
metatags
The metatags module contains functions for generating meta tags. Meta tag generation and extraction module.
tokiotokio
Async file helpers on the Tokio runtime (feature tokio). Async file helpers on the Tokio runtime (feature tokio).
utils
The utils module contains utility functions for metadata processing. Utility functions for metadata processing and HTML manipulation.

Functions§

extract_and_prepare_metadata
Extracts metadata from the content, generates keywords based on the metadata, and prepares meta tag groups.
extract_keywords
Extracts keywords from the metadata.

Type Aliases§

Keywords
Type alias for a list of keywords.
MetadataMapstd
Type alias for a map of metadata key-value pairs.
MetadataResult
Type alias for the result of metadata extraction and processing.