Public record
Software health reportschema 0.34.0 · metrics 2.10.0 · 2026-08-22 22:30 UTC

firecrawl / pdf-inspector

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

RustMIT★ 16,490 stars⑂ 1,143 forkssince Feb 2026View on GitHub ↗
KindLibraryCommand-line toolhow this is determined

firecrawl/pdf-inspector holds a health index of 86 out of 100, placing it in the Excellent band. It scores highest on Vitality (89/100) and lowest on Sustainability & Governance (56/100). It was last updated today. A single contributor accounts for most of its recent work.

86
overall / 100
Excellent

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

86
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 72 is calibrated to 86 on the published index scale (record calibration 2026-08-02).

Ownership

FirecrawlOrganization
3,430 followers109 public repossince May 2023

This repository is backed by an organization — shared, accountable stewardship that can outlive any single maintainer.

Package ecosystems

RegistryPackageVersionDownloads / moVersionsLast publishTags
crates.iopdf-inspector1.17.059,401151 day ago
PyPIpdf-inspector1.17.0318,242161 day ago
npm@firecrawl/pdf-inspector1.17.0228,885581 day agopdfpdf-extractionpdf-parsertext-extractionocrpdf-classificationnapirustfirecrawl

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

89Excellent · 21% of overall
How it's scored
36/36Push recencylast push 0 days ago
17.3/36Commit cadence25/52 weeks with commits
18/18Commit volume491 commits in the last year
10/10OpenSSF Scorecard: Maintained30 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
Inputs used
commits_last_year491
human_commit_share1
days_since_last_push0
active_weeks_last_year25

Release discipline

100Exceptional
How it's scored
27/27Ships releases3 releases published
36/36Release recencylatest release 5 days ago
27/27Release cadencea release every ~3.4 days
0/10OpenSSF Scorecard: Signed-Releasesno data
Inputs used
releases_count3
latest_release_tagv1.15.0
releases_from_tagsno
days_since_latest_release5
mean_days_between_releases3.4
Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

79Good · 17% of overall
How it's scored
60/60Stars16,490 stars
25/25Forks1,143 forks
9.1/15Watchers45 watchers
Inputs used
forks1,143
stars16,490
watchers45
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (MIT)
0/18CONTRIBUTING guide
0/13.5Code of conduct
0/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges4
has_contributingno
has_issue_templateno
has_code_of_conductno
readme_badge_servicesshields.io
has_pull_request_templateno
How it's scored
77.1/80Monthly downloads606,528 downloads/month across crates, npm, pypi
0/20Registry dependentsnot reported by this ecosystem
Inputs used
packagespdf-inspector, pdf-inspector, @firecrawl/pdf-inspector
dependents
ecosystemscrates, npm, pypi
total_downloads178,202
monthly_downloads606,528
unverified_packages_excluded
Excluded from scoring (no data or not applicable): Registry dependents. Remaining weights renormalized.

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

56Moderate · 23% of overall
How it's scored
9/54Bus factor1 contributor(s) cover half of all commits
0.6/22.5Commit distributiontop contributor authored 97% of commits
13.5/13.5Contributor breadth12 contributors
3/10OpenSSF Scorecard: Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
Inputs used
bus_factor1
contributors_sampled12
top_contributor_share0.974
How it's scored
9.3/42Issue resolution22% of issues closed
28/30PR acceptance236/253 decided PRs merged
0/13Newcomer PR acceptance0/5 first-time contributors' PRs merged in 30d
0/15OpenSSF Scorecard: Code-ReviewFound 0/30 approved changesets -- score normalized to 0
Inputs used
merged_prs236
open_issues74
closed_issues21
prs_merged_7d37
prs_decided_7d41
prs_merged_30d55
prs_decided_30d60
issue_closed_ratio0.221
closed_unmerged_prs17
first_time_authors_30d5
first_time_prs_merged_30d0
first_time_prs_decided_30d5
How it's scored
30/30Ownership backingorganization-owned
0/20Verified domain
25/25Owner reach3,430 followers of firecrawl
19.5/25Track record109 public repos, account ~3 yr old
Inputs used
followers3,430
owner_typeOrganization
is_verifiedno
owner_loginfirecrawl
public_repos109
account_age_days1,180

Package maintenance

100Exceptional
How it's scored
25/25Published & resolvable3 package(s) on crates, npm, pypi
35/35Publish recencylatest publish 1 days ago
20/20Version history58 published versions
20/20Not deprecatedactive, not deprecated or yanked
Inputs used
packagespdf-inspector, pdf-inspector, @firecrawl/pdf-inspector
ecosystemscrates, npm, pypi
any_deprecatedno
min_days_since_publish1

Engineering Quality

Are baseline engineering and documentation practices in place?

77Good · 19% of overall
How it's scored
24/24CI workflows6 workflow(s)
24/24Tests present
0/16Linter config
0/9.6Pre-commit hooks
0/6.4.editorconfig
20/20OpenSSF Scorecard: CI-Tests30 out of 30 merged PRs checked by a CI test -- score normalized to 10
Inputs used
has_ciyes
has_testsyes
has_editorconfigno
has_linter_configno
has_precommit_configno

Documentation

90Excellent
How it's scored
30/30README
25/25Documentation directory
15/15Documentation / homepage sitehttps://firecrawl.github.io/pdf-inspector/
10/10Repository description
10/10Topics10 topics
0/10Wiki
Inputs used
topicsmarkdown, nodejs, pdf, pdf-extraction, pdf-parser, python, rust, text-extraction, ocr-routing, pdf-classification
has_wikino
homepagehttps://firecrawl.github.io/pdf-inspector/
docs_sitehttps://firecrawl.github.io/pdf-inspector/
has_readmeyes
has_docs_diryes
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

59Moderate · 16% of overall
How it's scored
7.5/7.5Binary-Artifactsno binaries found in the repo
0/7.5Branch-Protectionbranch protection not enabled on development/release branches
2.5/2.5CI-Tests30 out of 30 merged PRs checked by a CI test -- score normalized to 10
0/2.5CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0/7.5Code-ReviewFound 0/30 approved changesets -- score normalized to 0
0.8/2.5Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
10/10Dangerous-Workflowno dangerous workflow patterns detected
0/7.5Dependency-Update-Toolno update tool detected
0/5Fuzzingproject is not fuzzed
2.5/2.5Licenselicense file detected
7.5/7.5Maintained30 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
5/5Packagingpackaging workflow detected
4.5/5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 9
0/5SASTSAST tool is not run on all commits -- score normalized to 0
5/5Security-Policysecurity policy file detected
0/7.5Signed-Releasesno data
7.5/7.5Token-PermissionsGitHub workflow tokens follow principle of least privilege
4.5/7.5Vulnerabilities4 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated17
scorecard_versionv5.5.0
checks_inconclusive1
scorecard_aggregate5.9
Excluded from scoring (no data or not applicable): Signed-Releases. Remaining weights renormalized.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

83Excellent · 4% of overall
How it's scored
45/45Agent instructionsAGENTS.md, CLAUDE.md
0/15Machine-readable docs (llms.txt)
40/40Legible commit history100 of 100 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share1
agent_instruction_filesAGENTS.md, CLAUDE.md
agent_instruction_max_bytes4,966
How it's scored
12.6/18One-command bootstrapCargo.toml, napi/Cargo.toml, wasm/Cargo.toml (toolchain convention, no task runner)
22/22Automated tests
0/11Lint / format config
11/11Static type checkingRust (statically typed)
10/10Reproducible environmentlockfile
10/10Demonstrated agent practice23 of the last 100 commits agent-authored or agent-credited
0/8Automated maintenanceno automated dependency updates observed
9/10OpenSSF Scorecard: Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 9
Inputs used
has_nixno
has_testsyes
lockfilesCargo.lock
has_dockerfileno
typed_languageyes
bootstrap_files
has_devcontainerno
has_linter_configno
typecheck_configs
agent_commit_share0.23
toolchain_manifestsCargo.toml, napi/Cargo.toml, wasm/Cargo.toml
dependency_bot_commit_share0
How it's scored
45/45Type-checkable codeRust (statically typed)
40.7/55Manageable file sizes18/69 source files over 60KB
Inputs used
primary_languageRust
largest_source_bytes348,818
source_files_sampled69
oversized_source_files18
How it's scored
0/40API schema (OpenAPI/GraphQL/proto)not applicable to this kind of software
0/20MCP servernot applicable to this kind of software
40/40Runnable examplesexamples
Inputs used
example_dirsexamples
has_mcp_signalno
api_schema_files
interfaces_expected_of
Excluded from scoring (no data or not applicable): API schema (OpenAPI/GraphQL/proto), MCP server. Remaining weights renormalized.

Key facts

16,490GitHub stars
12contributors
491commits, last 12 months
0days since last push
3releases
1bus factor
74open issues
crates.io, npm, PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token
  • Could not fetch crates package 'pdf-inspector-napi' from its registry
  • Could not fetch crates package 'pdf-inspector-wasm' from its registry
  • GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

More detail

Star and fork history 0 ★ / 1,143 ⇿
0Stars
1,143Forks
3Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

Only the most recent history is shown — this repository exceeds the collection window, so the earliest history is not captured.

02004006008001,0001,2001,1431642026-072026-072026-08
Major 0Minor 1Patch 1
OpenSSF Scorecard 5.9 / 10
5.9aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-22 22:30 UTC

10Binary-Artifactsno binaries found in the repo
0Branch-Protectionbranch protection not enabled on development/release branches
10CI-Tests30 out of 30 merged PRs checked by a CI test -- score normalized to 10
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0Code-ReviewFound 0/30 approved changesets -- score normalized to 0
3Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
10Dangerous-Workflowno dangerous workflow patterns detected
0Dependency-Update-Toolno update tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
10Maintained30 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
10Packagingpackaging workflow detected
9Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 9
0SASTSAST tool is not run on all commits -- score normalized to 0
10Security-Policysecurity policy file detected
n/aSigned-Releasesno releases found
10Token-PermissionsGitHub workflow tokens follow principle of least privilege
6Vulnerabilities4 existing vulnerabilities detected
Direct dependencies 16
RegistryPackageVersion constraintManifest
crates.iopyo30.25Cargo.toml
crates.iothiserror2.0Cargo.toml
crates.iolog0.4Cargo.toml
crates.ioregex1.10Cargo.toml
crates.ioonce_cell1.19Cargo.toml
crates.iounicode-normalization0.1Cargo.toml
crates.iottf-parser0.25Cargo.toml
crates.iopdf-inspectornapi/Cargo.toml
crates.ionapi3.0.0napi/Cargo.toml
crates.ionapi-derive3.0.0napi/Cargo.toml
crates.ioconsole_error_panic_hook0.1wasm/Cargo.toml
crates.iojs-sys0.3wasm/Cargo.toml
crates.iopdf-inspectorwasm/Cargo.toml
crates.ioserde1wasm/Cargo.toml
crates.ioserde-wasm-bindgen0.6wasm/Cargo.toml
crates.iowasm-bindgen0.2wasm/Cargo.toml
All dependencies not collected

The resolved dependency set could not be collected for this report: GitHub dependency-graph SBOM unavailable (404); the dependency graph may be disabled for this repository

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.34.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticscrates.io, PyPI, npm.