An AI agent that evaluates biodiversity, ecology, and environmental science datasets for AI-readiness using the FAIR4AI-Bio checklist. The agent reads a dataset's landing page or local metadata files, rates each of the 96 checklist items as meets / partial / does not meet / N/A (per RATING_RUBRIC.md), and produces a structured JSON report with reproducible FAIR4AI scores across five dimensions (Findable, Accessible, Interoperable, Reusable, AI-ready) plus an overall score — each in 0–1, where 1 is "most FAIR4AI", computed by the fair4ai-scoring skill.
The core thesis: FAIR compliance is necessary but not sufficient for AI-ready data. This agent surfaces the gap.
See example_outputs/ for complete evaluation reports for several NEON datasets.
- A dataset source: URL to a landing page, or a local directory containing metadata files (JSON-LD, EML, DataCite XML, README, etc.)
- Claude Code installed
Open Claude Code with fair4ai-eval-agent/ as the working directory:
cd fair4ai-eval-agent
claudeNo additional setup steps are required. The /evaluate-dataset skill is defined in .claude/skills/evaluate-dataset/SKILL.md and is tracked in this repository.
In the Claude Code chat, type:
/evaluate-dataset
The agent will prompt you for the dataset source and other parameters. Press Enter to accept defaults for optional parameters.
To evaluate many datasets at once, use the /batch-evaluate-datasets skill:
/batch-evaluate-datasets
It reads a flexible dataset list — a .txt, .csv, or .md file — parses out each dataset and its URL regardless of layout, and normalizes it into a standard CSV (email,name,dataset_short_name,url,notes). It then coordinates one sub-agent per dataset in parallel (launched in waves of 5 by default), each running the /evaluate-dataset workflow non-interactively. The default sub-agent model is Claude Haiku 4.5 (claude-haiku-4-5-20251001); the skill asks whether to use it or a different model.
Every artifact of a run is written into one self-contained run folder (default batch_run_<date>/):
batch_run_<date>/
batch_datasets_<date>.csv # normalized reference list (from any input format)
batch_evaluate_progress.md # resumable progress tracker
confidence_map.json
evaluation_results/FAIR4AI_eval_*.json # one full 96-item evaluation per dataset
fair4ai_scores_summary_<date>.csv # compiled per-dataset scores + status counts
fair4ai_score_distributions_<date>.{png,pdf,svg} # summary figure
FAIR4AI_summary_report_<date>.md # cross-dataset narrative report
The run is resumable — re-invoking the skill on an existing run folder reads batch_evaluate_progress.md and re-runs only the datasets that are still pending or failed. The summary figure needs numpy + matplotlib (pip install -r scripts/requirements-viz.txt); if they're absent the figure is skipped and the CSV + report are still produced (scoring and compilation are stdlib-only).
Claude Code reads these files automatically when you start a session in this directory:
CLAUDE.md— project context loaded into every session: what the agent does, how the checklist is structured, the output JSON schema, and expected score patterns..claude/skills/evaluate-dataset/SKILL.md— the/evaluate-datasetskill. Contains the step-by-step evaluation workflow: gather parameters, load the checklist, fetch metadata, rate each item, build the output JSON, and compute scores.RATING_RUBRIC.md— the authority for how each item'smeets / partial / does not meet / N/Astatus is chosen (including the N/A rule)..claude/skills/fair4ai-scoring/SKILL.md+scripts/compute_fair4ai_scores.py— the skill and deterministic tool that computesummary.fair4ai_scores(0–1) from the per-item statuses and each item'sFAIR4AI category..claude/skills/batch-evaluate-datasets/SKILL.md— the/batch-evaluate-datasetsskill: reads a dataset list, coordinates parallel per-dataset sub-agents, and compiles a scores CSV, summary figure, and report into one run folder. It reusesevaluate-dataset(in Batch / non-interactive mode) plus thescripts/compile_fair4ai_results.pyandscripts/make_fair4ai_figure.pyhelpers.
Both capabilities live under .claude/skills/ (one skill per directory, each with a SKILL.md). To modify agent behavior, edit these files directly. Changes are tracked in git and shared across the team.
The JSON report has three top-level sections:
session— evaluation date, AI model, metadata sources used, dataset identity (title, DOI, landing page URL, citation), and evaluator informationresponses— one object per checklist item withsection,sub_section,question,status(meets/partial/does not meet/N/A),evidence,notes,recommendation, andfair4ai_category(the dimension(s) the item counts toward)summary—strengths,gaps,overall_assessment(2–3 sentence narrative), andfair4ai_scores(each dimension plus an overall score in 0–1, with per-dimensiondetailscounts) computed by thefair4ai-scoringskill
Output filename convention: FAIR4AI_eval_<dataset-name>_<YYYY-MM-DD>.json
See example_outputs/ for complete examples.
| File | Description |
|---|---|
CLAUDE.md |
Project context auto-loaded by Claude Code each session |
.claude/skills/evaluate-dataset/SKILL.md |
The /evaluate-dataset skill and evaluation workflow (supports Batch / non-interactive mode) |
.claude/skills/fair4ai-scoring/SKILL.md |
The /fair4ai-scoring skill wrapping the scoring script |
.claude/skills/batch-evaluate-datasets/SKILL.md |
The /batch-evaluate-datasets skill: parallel batch evaluation → scores CSV + figure + report |
scripts/compute_fair4ai_scores.py |
Deterministic 0–1 FAIR4AI scoring tool (stdlib-only) |
scripts/compile_fair4ai_results.py |
Compiles a folder of evaluation JSONs into a scores CSV + aggregates (stdlib-only) |
scripts/make_fair4ai_figure.py |
Renders the 6-panel score-distribution figure (needs numpy + matplotlib) |
scripts/requirements-viz.txt |
Optional deps (numpy, matplotlib) for the figure only |
example_inputs/ |
Sample batch input lists (.csv and .md) for /batch-evaluate-datasets |
CHECKLIST.csv |
96-item FAIR4AI-Bio checklist (8 sections; each item mapped to EML, DataCite, Schema.org, Croissant, and its FAIR4AI dimension(s)) |
RATING_RUBRIC.md |
Authority for choosing each item's meets / partial / does not meet / N/A status |
CHECKLIST_OVERVIEW.md |
Narrative description of all 8 checklist sections |
QUICKSTART.md |
Step-by-step usage guide with example sessions |
example_outputs/ |
Complete evaluation reports for NEON datasets (example outputs) |
- General Information — bibliographic metadata, data dictionary, machine-readiness flag
- Data Structure — file formats, dataset organization, variables, technical specs
- Source Data — collection type, instrumentation/sensor metadata, sampling design
- Data Processing — transformations, gap-filling, annotation provenance, train/val/test split definitions
- Data Quality — completeness, consistency, integrity, timeliness
- Guidance & Recommendations — biases, class imbalance, non-detections, prior AI/ML usage history
- Data Access — delivery options, format openness, license (SPDX identifier), privacy controls
- Provenance — citation (DOI, ORCIDs, checksums), processing platform, CARE Principles and ethical governance