Public record
Software health reportschema 0.31.0 · metrics 2.10.0 · 2026-08-12 19:48 UTC

google / langextract

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

PythonApache-2.0★ 38,312 stars⑂ 2,683 forkssince Jul 2025View on GitHub ↗
KindPluginLibraryhow this is determined

google/langextract holds a health index of 93 out of 100, placing it in the Exceptional band. It scores highest on Engineering Quality (92/100) and lowest on AI Readiness (61/100). It was last updated 1 day ago. A single contributor accounts for most of its recent work.

93
overall / 100
Exceptional

Software health index

Metrics are grouped into weighted categories on one standardized 1–100 scale. Overall starts as their weighted mean, calibrated against the distribution of the public record so bands carry percentile meaning; when public evidence triggers the High-Risk Jurisdiction Policy, the rating is adjusted and receives an At Risk ceiling of 34.

93
Exceptional93-100The record's top tier (≈ top 5%); essentially all checked criteria met
Excellent80-92Strong across the board; minor gaps
Good65-79Healthy; gaps are limited and manageable
Moderate50-64Acceptable with notable gaps; review recommended
Weak35-49Material weaknesses across several areas
At Risk20-34Significant weaknesses; adoption warrants caution
Critical1-19Severe problems (abandoned, single-maintainer, no hygiene)
VitalityCommunity &AdoptionSustainability &GovernanceEngineeringQualitySecurityAI Readiness

Score profile

Each axis is a category. The shape matters more than the average — a healthy subject fills the whole shape, while a spike-and-crater profile means strength in one dimension is masking risk in another.

The weighted overall 79 is calibrated to 93 on the published index scale (record calibration 2026-08-02).

Ownership

GoogleOrganization
76,871 followers2,891 public repossince Jan 2012

This repository is backed by an organization — shared, accountable stewardship that can outlive any single maintainer.

Package ecosystems

RegistryPackageVersionDownloads / moVersionsLast publish
PyPIlangextract1.6.0-2441 days ago

Metrics by category

Vitality

Is the project alive — is code being written and are releases shipping?

89Excellent · 21% of overall
How it's scored
36/36Push recencylast push 1 days ago
16.6/36Commit cadence24/52 weeks with commits
18/18Commit volume113 commits in the last year
10/10OpenSSF Scorecard: Maintained18 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
Inputs used
commits_last_year113
human_commit_share1
days_since_last_push1
active_weeks_last_year24

Release discipline

100Exceptional
How it's scored
27/27Ships releases18 releases published
36/36Release recencylatest release 41 days ago
27/27Release cadencea release every ~35.7 days
0/10OpenSSF Scorecard: Signed-Releasesno data
Inputs used
releases_count18
latest_release_tagv1.6.0
releases_from_tagsno
days_since_latest_release41
mean_days_between_releases35.7
Excluded from scoring (no data or not applicable): OpenSSF Scorecard: Signed-Releases. Remaining weights renormalized.

Community & Adoption

Does the project have users, downloads, attention, and a welcoming setup for contributors?

91Excellent · 17% of overall
How it's scored
60/60Stars38,312 stars
25/25Forks2,683 forks
12.4/15Watchers169 watchers
Inputs used
forks2,683
stars38,312
watchers169
growth_stateunverified
growth_factor_pct100
growth_unverified_reasonno_history

Community health

85Excellent
How it's scored
22.5/22.5README
22.5/22.5Licenserecognized license (Apache-2.0)
18/18CONTRIBUTING guide
13.5/13.5Code of conduct
0/7.2Issue template
0/6.3PR template
Inputs used
has_readmeyes
has_licenseyes
readme_badges4
has_contributingyes
has_issue_templateno
has_code_of_conductyes
readme_badge_servicesgithub.com, shields.io
has_pull_request_templateno

Sustainability & Governance

Will the project survive its people — bus factor, responsiveness, who backs it, and package upkeep?

65Good · 23% of overall
How it's scored
9/54Bus factor1 contributor(s) cover half of all commits
3.2/22.5Commit distributiontop contributor authored 86% of commits
13.5/13.5Contributor breadth22 contributors
3/10OpenSSF Scorecard: Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
Inputs used
bus_factor1
contributors_sampled22
top_contributor_share0.86
How it's scored
26.1/42Issue resolution62% of issues closed
16.1/30PR acceptance120/224 decided PRs merged
0/13Newcomer PR acceptance0/2 first-time contributors' PRs merged in 30d
1.5/15OpenSSF Scorecard: Code-ReviewFound 3/30 approved changesets -- score normalized to 1
Inputs used
merged_prs120
open_issues74
closed_issues121
prs_merged_7d1
prs_decided_7d1
prs_merged_30d8
prs_decided_30d10
issue_closed_ratio0.621
closed_unmerged_prs104
first_time_authors_30d2
first_time_prs_merged_30d0
first_time_prs_decided_30d2
How it's scored
30/30Ownership backingorganization-owned
0/20Verified domainverified-domain status not read for this organization
25/25Owner reach76,871 followers of google
25/25Track record2,891 public repos, account ~14 yr old
Inputs used
followers76,871
owner_typeOrganization
is_verified
owner_logingoogle
public_repos2,891
account_age_days5,320
Excluded from scoring (no data or not applicable): Verified domain. Remaining weights renormalized.

Package maintenance

100Exceptional
How it's scored
25/25Published & resolvable1 package(s) on pypi
35/35Publish recencylatest publish 41 days ago
20/20Version history24 published versions
20/20Not deprecatedactive, not deprecated or yanked
Inputs used
packageslangextract
ecosystemspypi
any_deprecatedno
min_days_since_publish41

Engineering Quality

Are baseline engineering and documentation practices in place?

92Excellent · 19% of overall
How it's scored
24/24CI workflows11 workflow(s)
24/24Tests present
16/16Linter config.pylintrc, tox.ini
9.6/9.6Pre-commit hooks
0/6.4.editorconfig
20/20OpenSSF Scorecard: CI-Tests29 out of 29 merged PRs checked by a CI test -- score normalized to 10
Inputs used
has_ciyes
has_testsyes
has_editorconfigno
has_linter_configyes
has_precommit_configyes

Documentation

90Excellent
How it's scored
30/30README
25/25Documentation directory
15/15Documentation / homepage sitehttps://pypi.org/project/langextract/
10/10Repository description
10/10Topics11 topics
0/10Wiki
Inputs used
topicsllm, nlp, python, gemini-ai, information-extration, large-language-models, structured-data, gemini, gemini-api, gemini-flash, gemini-pro
has_wikino
homepagehttps://pypi.org/project/langextract/
docs_sitehttps://pypi.org/project/langextract/
has_readmeyes
has_docs_diryes
has_descriptionyes

Security

Are visible security and supply-chain practices strong, without unresolved high-risk jurisdiction exposure?

63Moderate · 16% of overall
How it's scored
7.5/7.5Binary-Artifactsno binaries found in the repo
6/7.5Branch-Protectionbranch protection is not maximal on development and all release branches
2.5/2.5CI-Tests29 out of 29 merged PRs checked by a CI test -- score normalized to 10
0/2.5CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
0.8/7.5Code-ReviewFound 3/30 approved changesets -- score normalized to 1
0.8/2.5Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
0/10Dangerous-Workflowdangerous workflow patterns detected
7.5/7.5Dependency-Update-Toolupdate tool detected
0/5Fuzzingproject is not fuzzed
2.5/2.5Licenselicense file detected
7.5/7.5Maintained18 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
5/5Packagingpackaging workflow detected
0/5Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0/5SASTSAST tool is not run on all commits -- score normalized to 0
5/5Security-Policysecurity policy file detected
0/7.5Signed-Releasesno data
0/7.5Token-Permissionsdetected GitHub workflow tokens with excessive permissions
7.5/7.5Vulnerabilities0 existing vulnerabilities detected
Inputs used
sourceopenssf_scorecard
checks_evaluated17
scorecard_versionv5.5.0
checks_inconclusive1
scorecard_aggregate5.4
Excluded from scoring (no data or not applicable): Signed-Releases. Remaining weights renormalized.

Dependency advisories

100Exceptional
How it's scored
35/35Direct dependencies free of known advisoriesno direct dependency carries a known advisory
25/25Indirect dependencies free of known advisoriesno indirect dependency carries a known advisory
0/40No advisories left outstandingno advisory carries a publication date
Inputs used
sourceosv
advisories0
affected_packages0
assessed_packages54
unassessed_packages0
affected_by_severitynone
direct_affected_packages0
Excluded from scoring (no data or not applicable): No advisories left outstanding. Remaining weights renormalized. Matched the pypi:langextract@1.6.0 runtime dependency closure — what installing the published package pulls in — 54 packages. Reachability is not analyzed.

AI Readiness

How well is the repo equipped to be developed and maintained with AI coding agents? Carries a deliberately small weight (4%): agent tooling is a real maintenance signal, but a repository with none can still reach 100/100.

61Moderate · 4% of overall
How it's scored
0/45Agent instructionsno CLAUDE.md / AGENTS.md / editor rules
0/15Machine-readable docs (llms.txt)
40/40Legible commit history94 of 100 human commits state their intent (structured subject or explanatory body)
Inputs used
has_llms_txtno
llms_txt_url
legible_history_share0.94
agent_instruction_files
agent_instruction_max_bytes
How it's scored
0/18One-command bootstrap
22/22Automated tests
11/11Lint / format config.pylintrc, tox.ini
11/11Static type checkinglangextract/py.typed
10/10Reproducible environmentDockerfile
0/10Demonstrated agent practiceno agent-authored commits among the last 100
0/8Automated maintenanceno automated dependency updates observed
0/10OpenSSF Scorecard: Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
Inputs used
has_nixno
has_testsyes
lockfiles
has_dockerfileyes
typed_languageno
bootstrap_files
has_devcontainerno
has_linter_configyes
typecheck_configslangextract/py.typed
agent_commit_share0
toolchain_manifests
dependency_bot_commit_share0
How it's scored
27/45Type-checkable codePython with type-check config (langextract/py.typed)
53.8/55Manageable file sizes2/91 source files over 60KB
Inputs used
primary_languagePython
largest_source_bytes91,722
source_files_sampled91
oversized_source_files2
How it's scored
0/40API schema (OpenAPI/GraphQL/proto)not applicable to this kind of software
0/20MCP servernot applicable to this kind of software
40/40Runnable examplesexamples, notebooks
Inputs used
example_dirsexamples, notebooks
has_mcp_signalno
api_schema_files
interfaces_expected_of
Excluded from scoring (no data or not applicable): API schema (OpenAPI/GraphQL/proto), MCP server. Remaining weights renormalized.

Key facts

38,312GitHub stars
22contributors
113commits, last 12 months
1days since last push
18releases
1bus factor
74open issues
PyPIpackage ecosystems

Data collection warnings

  • Star history unavailable: GitHub GraphQL error: Resource not accessible by personal access token

More detail

Star and fork history 0 ★ / 2,683 ⇿
0Stars
2,683Forks
6Releases

When each star and fork was added, collected from GitHub and bucketed by day. Cumulative growth sits directly above the daily additions it is made of, so the two read against each other: steady organic accretion looks nothing like an abrupt, short-lived burst. Where that difference is measurable, it is reported as growth authenticity.

Only the most recent history is shown — this repository exceeds the collection window, so the earliest history is not captured.

1,5001,7502,0002,2502,5002,7502,6831202026-022026-052026-08
Major 0Minor 5Patch 1
OpenSSF Scorecard 5.4 / 10
5.4aggregate

Independent, tool-agnostic security assessment from the open-source OpenSSF Scorecard. Each check rewards a security practice, not a specific vendor's tool. Checks Scorecard could not determine are marked n/a and excluded from the security score (never counted as zero).Scorecard v5.5.0 · 2026-08-12 19:47 UTC

10Binary-Artifactsno binaries found in the repo
8Branch-Protectionbranch protection is not maximal on development and all release branches
10CI-Tests29 out of 29 merged PRs checked by a CI test -- score normalized to 10
0CII-Best-Practicesno effort to earn an OpenSSF best practices badge detected
1Code-ReviewFound 3/30 approved changesets -- score normalized to 1
3Contributorsproject has 1 contributing companies or organizations -- score normalized to 3
0Dangerous-Workflowdangerous workflow patterns detected
10Dependency-Update-Toolupdate tool detected
0Fuzzingproject is not fuzzed
10Licenselicense file detected
10Maintained18 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 10
10Packagingpackaging workflow detected
0Pinned-Dependenciesdependency not pinned by hash detected -- score normalized to 0
0SASTSAST tool is not run on all commits -- score normalized to 0
10Security-Policysecurity policy file detected
n/aSigned-Releasesno releases found
0Token-Permissionsdetected GitHub workflow tokens with excessive permissions
10Vulnerabilities0 existing vulnerabilities detected
Direct dependencies 17
RegistryPackageVersion constraintManifest
PyPIabsl-py>=1.0.0pyproject.toml
PyPIaiohttp>=3.8.0pyproject.toml
PyPIasync_timeout>=4.0.0pyproject.toml
PyPIexceptiongroup>=1.1.0pyproject.toml
PyPIgoogle-genai>=1.39.0pyproject.toml
PyPIgoogle-cloud-storage>=2.14.0pyproject.toml
PyPIml-collections>=0.1.0pyproject.toml
PyPImore-itertools>=8.0.0pyproject.toml
PyPInumpy>=1.20.0pyproject.toml
PyPIpandas>=1.3.0pyproject.toml
PyPIpydantic>=1.8.0pyproject.toml
PyPIpython-dotenv>=0.19.0pyproject.toml
PyPIPyYAML>=6.0pyproject.toml
PyPIregex>=2023.0.0pyproject.toml
PyPIrequests>=2.25.0pyproject.toml
PyPItqdm>=4.64.0pyproject.toml
PyPItyping-extensions>=4.0.0pyproject.toml
All dependencies 31

Full resolved dependency set from the GitHub dependency graph: 17 direct and 14 indirect (transitive) packages. The transitive closure is complete when the repository commits a lockfile.

RegistryPackageVersionRelation
PyPIabsl-pydirect
PyPIaiohttpdirect
PyPIasync-timeoutdirect
PyPIexceptiongroupdirect
PyPIgoogle-cloud-storagedirect
PyPIgoogle-genaidirect
PyPIml-collectionsdirect
PyPImore-itertoolsdirect
PyPInumpydirect
PyPIpandasdirect
PyPIpydanticdirect
PyPIpython-dotenvdirect
PyPIpyyamldirect
PyPIregexdirect
PyPIrequestsdirect
PyPItqdmdirect
PyPItyping-extensionsdirect
PyPIimport-linterindirect
PyPIipythonindirect
PyPIisort5.13.2indirect
PyPInotebookindirect
PyPIopenaiindirect
PyPIpre-commitindirect
PyPIpyink24.3.0indirect
PyPIpylintindirect
PyPIpytestindirect
PyPIpytypeindirect
PyPIsetuptoolsindirect
PyPItomliindirect
PyPItoxindirect
PyPItypes-regexindirect
Dependency advisories 0

Installing pypi:langextract@1.6.0 pulls in 54 packages, direct and transitive: 0 carry known advisories, of which 0 are direct dependencies.

No known advisories affect the assessed dependencies.

An advisory means the version recorded in the dependency graph falls inside an advisory’s affected range. Reachability is not analysed, and the graph includes development and test pins — a finding may concern tooling rather than shipped software.

Raw JSON report machine-readable

Feedback

Spotted something off in this report, or have thoughts to share? Wrong measurements, missed tooling, ideas, questions — anything is welcome. Every message is read and gets a response.

The message is kept through sign-in.

Scores are signals, not warranties. They reflect publicly visible practices on GitHub — not a code audit, and not a security guarantee.

Missing data is excluded and weights renormalized, never scored as zero. Methodology is versioned and open: metrics v2.10.0, schema v0.31.0 — full methodology · metrics wiki.

How one result sits in the wider record: aggregate statisticsPyPI.