Skip to content

Register governed post-Phase-2 research sources - #4

Open
EmergentMonk wants to merge 210 commits into
mainfrom
research-corpus-governed-expansion
Open

Register governed post-Phase-2 research sources#4
EmergentMonk wants to merge 210 commits into
mainfrom
research-corpus-governed-expansion

Conversation

@EmergentMonk

@EmergentMonk EmergentMonk commented Sep 2, 2026

Copy link
Copy Markdown
Member

Implements ROADMAP Workstream A: research-corpus expansion and source governance, while establishing bounded follow-up methodology for Workstreams F, G, H, and I.

Scope

This PR formally registers a governed post-Phase-2 source batch in docs/RESEARCH-REFERENCE-CORPUS.md rather than treating planning reports, media references, archival records, glossaries, or community discussions as adopted evidence by default.

Every adopted entry records attributable source destinations, source type and relevance, rights/provenance boundaries, epistemic limits, exactly one community-governance classification, research/project mappings, and a safe abstraction path for independently authored or appropriately licensed benchmark work.

Governed source batch

The registered batch covers:

  • Black Comedy
  • Kath & Kim
  • The Castle
  • Shaun Micallef's MAD AS HELL
  • Acropolis Now
  • Chey (2021)
  • Hurley (2025)
  • Slade, Australian Sketch Comedy Field Theory (ASCFT)
  • Trans-Tasman constitutional/federation context
  • ABC Language history of Australian sexual slang
  • Victoria University Australian slang dictionary
  • r/australia slang community-attestation material
  • Australian Defence multinational communication reports
  • WWII American-serviceman Australia language guides and archival records

Media sources remain mechanism-discovery references, not representative speech corpora. Scholarly, institutional, archival, theoretical, glossary, and community-attestation sources retain source-specific evidential and rights limits.

Roadmap and canonical methodology

The roadmap contains:

  • Workstream F, ASCFT-derived adversarial pragmatics
  • Workstream G, Trans-Tasman relational pragmatics and lexical context
  • Workstream H, Slang density, register compression, and operational intelligibility
  • Workstream I, Australian and United States policing-context transfer

docs/METHODOLOGY.md centrally defines all four experiment families. Workstream I is explicitly source-gated and is not legal advice or a ranking of which country polices “better.” Every implemented Workstream I item must record, at minimum, country, jurisdiction, agency or institutional role, encounter type, source date or version, registered source identifiers or links supporting any legal or procedural condition supplied to the model, and claim type.

Before publishing any family involving coercion, consent, search, detention, questioning, force, emergency powers, or legal rights, the governing sources must be verified as current for the recorded jurisdiction and date and the project must obtain appropriate Australian and United States legal, policing, civil-liberties, and community review.

A claim that Australian policing uses a lighter conversational touch is eligible only as a source-gated, encounter-specific hypothesis about register, discretion, or interactional style. It cannot become a universal description, a national moral ranking, or evidence that calm language removes coercive authority.

Candidate policing-context distinctions include:

  • US POLICE SCRIPT != AUSTRALIAN LEGAL PROCEDURE
  • POLICE TERMINOLOGY != CROSS-JURISDICTION EQUIVALENCE
  • CASUAL ADDRESS != FRIENDSHIP OR CONSENT
  • CALM TONE != ABSENCE OF COERCIVE AUTHORITY
  • POLITE WORDING != VOLUNTARY CHOICE
  • FICTIONAL POLICE TROPE != OPERATIONAL POLICY
  • ONE AGENCY != A NATIONAL POLICING SYSTEM
  • ONE ENCOUNTER != SYSTEM-WIDE GROUND TRUTH
  • JURISDICTIONAL DIFFERENCE != NATIONAL MORAL CHARACTER
  • LEGAL INFORMATION != LEGAL ADVICE

Latest Codex repair pass

The current exact-head hardening adds five further browser/rendered-semantics protections:

  • applies inline CSS cascade semantics to display and visibility, including declaration order and !important precedence, so a later visible declaration cannot be incorrectly treated as hidden because of an earlier display:none;
  • fails closed on visually hidden raw HTML tables inside governed entries rather than pretending the lightweight validator implements full HTML5 foster-parenting semantics;
  • pins the complete browser-visible Workstream H section with a SHA-256 fixture, so a separate visible companion statement cannot reverse the listener-identity safeguard while retaining the original affirmative sentence;
  • pins the complete browser-visible canonical policing methodology section with a SHA-256 fixture before metadata and high-stakes publication/review receipts, so companion prose cannot silently reverse those gates;
  • treats aria-hidden=true correctly as accessibility-tree metadata rather than visual hiding, keeping visually rendered ARIA-hidden content inside browser-visible integrity and uniqueness checks.

Cumulative hardening retained by this PR includes:

  • ordinary read-only .github/workflows/ci.yml; temporary repair runners self-delete after verified repair commits;
  • rendered registration-contract, Status, governed-batch, Workstream H, Trans-Tasman methodology, Workstream I, policing methodology, and per-entry section scoping with preserved offsets;
  • browser-hidden/non-graphical handling for the hidden attribute, CSS-hidden content, script, style, template, default-closed <details> and <dialog>, SVG <title>/<desc> metadata, self-closing non-void browser semantics, HTML5 implied-end behavior, and CSS-comment normalization; aria-hidden=true remains visually visible for integrity/uniqueness checks because it affects accessibility-tree exposure rather than visual rendering;
  • CommonMark-aware comments, code spans, fences, indentation, tabs, list/quote container ownership, nested containers, thematic breaks, document-scoped/multiline/container-scoped link-reference definitions, raw HTML blocks, paragraph continuations, balanced labels, escaped/entity destinations, multiline titles, and additional raw-HTML block families;
  • exact metadata-field uniqueness across equivalent Markdown emphasis, nested emphasis, raw HTML <strong>/<b>, nested/wrapped HTML markup, entity-encoded labels, Markdown-link-wrapped labels, and raw-HTML-anchor-wrapped labels;
  • exact registered-source destination sets covering inline links, full/collapsed/shortcut references, document-scoped and multiline reference definitions, autolinks, bare HTTPS URLs, raw HTML anchors, HTML entities, Markdown escapes, and multiline inline-link titles;
  • fail-closed rejection of rendered non-HTTPS/unusable provenance links, image-only sources, localhost/single-label hosts, private/loopback/non-global IPs, legacy numeric loopback spellings, special-use host suffixes, malformed destinations, interactive form controls, and visually hidden raw-table constructs in governed entries;
  • complete normalized SHA-256 integrity fixtures for source type, governance rationale, rights/provenance, epistemic status, safe abstraction, research mappings, project mappings, the registration contract, complete browser-visible governed-entry bodies, Workstream H, Workstream I, and the canonical policing methodology section;
  • DOI visible-value and navigable-destination pinning;
  • browser-visible Workstream H listener-variable, stereotype, and community-attestation safeguards;
  • browser-visible Workstream I source gates, all ten policing invariants, complete seven-field metadata minimum, affirmative full-line enforcement, and the canonical high-stakes publication/review gate;
  • the categorical stereotype boundary: sources may document that a stereotype existed, but exact group-stereotyping wording remains excluded from repository content and redistributable benchmark items;
  • the Phase 2 analysis boundary that free-text pragmatic interpretations remain qualitative evidence and are not assigned a misleading exact-string IAA score;
  • the complete pre-existing Phase 1 changelog history.

Safeguards

  • RESEARCH REFERENCE != REDISTRIBUTABLE DATA
  • No copied dialogue, scripts, subtitles, transcripts, clips, catchphrases, body-camera audio, identifiable encounter material, or other third-party expression is added as benchmark data.
  • No community-specific benchmark ground truth is invented from media or reference material.
  • No new humour-mechanism labels are promoted by this PR.
  • FORMAL ANALOGY != PHYSICAL ONTOLOGY
  • MATHEMATICAL MODEL != EMPIRICALLY VALIDATED MECHANISM
  • DEADPAN DELIVERY != LITERAL INTENT

Verification

CI #635 is green on exact head 75dd00bbe6a545ad3cc5b40641874366c68bceab using the ordinary read-only workflow:

  • Python 3.11: pass
  • Python 3.12: pass
  • wheel/schema smoke test: pass
  • starter dataset: 15 records valid
  • Phase 2 pilot pack: 60 items valid
  • focused governance suite: 194 tests passed
  • full suite: 340 tests passed

All currently known Codex review threads are resolved.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @EmergentMonk, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 1 day and 19 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 2, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-09-03T21:03:21.610282Z 75dd00b Manual request
🔒 Security Review Completed 2026-09-02T15:47:40.153522Z 42c6e7d PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@sourcery-ai

sourcery-ai Bot commented Sep 2, 2026

Copy link
Copy Markdown

Reviewer's Guide

This PR establishes a governed process for expanding the post-Phase-2 research corpus, registers five media references and two scholarly sources with explicit epistemic and rights boundaries, and adds regression checks to prevent redistribution or unsupported community-specific benchmark claims.

Sequence diagram for safe research-source registration

sequenceDiagram
    participant Researcher
    participant Registry as ResearchReferenceRegistry
    participant Tests as RegressionChecks
    participant Benchmark as BenchmarkDesign

    Researcher->>Registry: Register source metadata
    Registry-->>Researcher: Record rights, provenance, epistemic status, mappings
    Researcher->>Benchmark: Derive safe abstraction
    Benchmark-->>Researcher: Use original or appropriately licensed work
    Researcher->>Tests: Validate registry boundaries
    Tests-->>Researcher: Enforce non-redistribution and consultation gate
    Tests-->>Researcher: Reject unsupported community-specific ground truth
Loading

Flow diagram for governed research-source registration

flowchart LR
    A[Proposed research source] --> B[Registration contract]
    B --> C[Source type and rationale]
    B --> D[Rights and provenance boundary]
    B --> E[Explicit epistemic status]
    B --> F[Research mappings and invariants]
    B --> G[Safe benchmark abstraction]
    G --> H[Independently authored or appropriately licensed work]
    D --> I[No copyrighted expression redistributed]
    E --> J[No unsupported community-specific ground truth]
Loading

File-Level Changes

Change Details Files
Introduces a formal governance contract and registers a bounded post-Phase-2 research-source batch without treating references as adopted data or taxonomy changes.
  • Defines required source links, type, rationale, rights/provenance, epistemic status, research mappings, invariants, and safe-abstraction guidance.
  • Registers five media references and two peer-reviewed studies with source-specific research mappings and limitations.
  • Preserves explicit non-redistribution, non-representativeness, consultation, and community-specific ground-truth safeguards.
  • Adds political-satire and historical/community-context boundaries for interpretation.
docs/RESEARCH-REFERENCE-CORPUS.md
Adds regression coverage for the registration contract and governance invariants.
  • Checks the registration sections and all seven registered entries are present.
  • Checks rights, epistemic-status, safe-abstraction, consultation, provenance, permissions, scope, and non-redistribution language.
tests/test_research_reference_registry.py
Documents the governed corpus expansion in the changelog.
  • Records the new source-registration contract and first governed media and scholarship batch.
  • Records the regression coverage for governance boundaries.
CHANGELOG.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

Copy link
Copy Markdown
Member Author

@codex review

Please review exact current head 42c6e7df7b03a54a6483858c1bc7e111f1212931.

Scope is the first governed implementation of ROADMAP Workstream A:

  • formal research-source registration contract;
  • seven post-Phase-2 registered media/scholarly references;
  • explicit rights/provenance and epistemic-status boundaries;
  • safe benchmark-abstraction rules;
  • stronger consultation gate for community-specific material;
  • no benchmark dialogue/data and no taxonomy promotion;
  • focused regression guarding the registration/governance contract.

CI run #250 is green on this exact head with 146 tests passing on Python 3.11 and 3.12, 15 starter records valid, 60 pilot items valid, and wheel/schema smoke passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 42c6e7df7b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/RESEARCH-REFERENCE-CORPUS.md Outdated
Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

Please review exact head 1527ac75057bd037a3ee1c1fae93e1bafbb03e10, with particular attention to the mandatory per-source registration contract, per-entry governance regression coverage, and community-specific consultation boundaries.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1527ac7505

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated
Comment thread docs/RESEARCH-REFERENCE-CORPUS.md Outdated
Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

The three findings from review of 1527ac75057bd037a3ee1c1fae93e1bafbb03e10 are addressed on current head ae7b0c595213c5016da8ed80ebf525994d5f44d5.

Please re-check specifically:

  • the governed post-Phase-2 batch is discovered dynamically so newly added ### entries cannot bypass validation;
  • every discovered entry must contain non-empty mandatory governance values, non-empty research/project mappings, and at least one HTTPS registered-source link;
  • The Castle now carries the class-marked consultation/provenance/permissions/scope-limitation boundary and is regression-checked as a community-specific/class-marked source.

CI run #258 is green on this exact head with 146 tests passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ae7b0c5952

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py
Comment thread tests/test_research_reference_registry.py Outdated
Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

Please review exact head e93118be41a78d874737581bdcaa85bcffa8ef48. The latest fixes isolate registered-source URLs from DOI/later metadata, encode per-entry community-specific governance classification with automatic consultation-boundary enforcement, and reject duplicate registered-entry headings before section construction. CI #262 is green with 146 tests passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e93118be41

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated
Comment thread tests/test_research_reference_registry.py Outdated
Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

Please review exact head 24f8f04f8bbc7c239954fcc50841aa2c063b6070. The latest three validator findings are addressed by same-line scalar metadata parsing, same-line community-governance rationale parsing, and enforcing the consultation boundary specifically inside Safe benchmark abstraction. Three focused mutation regressions cover those cases. CI #264 is green on this exact head with 149 tests passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 24f8f04f8b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

Please review exact head 6a94058b2a1d8ef81c9d5008917835b4afb553c2 after the latest registry-validator hardening. In particular, verify that each governed entry now permits exactly one Community-specific governance classification and that contradictory duplicate classifications fail closed. CI #266 is green with 150 tests passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6a94058b2a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated

Copy link
Copy Markdown
Member Author

@codex review

Please re-review exact head b82c2db4293fb4af8f20657ce3884393443c2ce0. The final outstanding validator issue has been addressed: all community-governance field occurrences are counted before value parsing, exactly one field is required, and malformed duplicate fields now fail closed. CI #268 is green with 151 tests passing.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b82c2db429

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated
@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex please review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0450a99684

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated
Comment thread tests/test_research_reference_registry.py
Comment thread tests/test_policing_context_roadmap.py
Comment thread tests/test_research_reference_registry.py
Comment thread tests/test_research_reference_registry.py
Comment thread tests/test_research_reference_registry.py
@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex please review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c52ad8230a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/test_research_reference_registry.py Outdated
Comment thread tests/test_research_reference_registry.py
Comment thread tests/test_workstream_h_methodology.py
Comment thread tests/test_policing_contract_receipt.py
Comment thread tests/test_research_reference_registry.py Outdated
@EmergentMonk

Copy link
Copy Markdown
Member Author

@codex please review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 75dd00bbe6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +681 to +684
def _is_hidden(tag: str, attrs: list[tuple[str, str | None]]) -> bool:
tag = tag.lower()
if tag in {"script", "style", "template"}:
return True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat ordinary HTML title elements as non-rendering

Beyond the SVG-title repair, wrapping a pinned rights clause in <title>...</title> still lets _validate_registry_corpus() pass because _is_hidden() hides title only when it is under svg. An ordinary HTML title element does not render visible body text, so this can remove a source's provenance or redistribution boundary from the document while preserving every field and integrity hash; classify HTML metadata title content as non-visible or reject it in governed entries.

AGENTS.md reference: AGENTS.md:L98-L105

Useful? React with 👍 / 👎.

Comment on lines +633 to +635
def _css_hides_element(style: str) -> bool:
"""Apply inline CSS declaration order and !important precedence."""
cleaned = CSS_COMMENT_PATTERN.sub("", style.lower())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Decode CSS escapes before evaluating visibility

Beyond the CSS-comment and cascade repairs, wrapping a required clause in <span style="display:n\6f ne">...</span> still passes full registry validation: CSS decodes \6f to o, so the browser applies display:none, while this literal comparison treats the clause as visible and leaves all pinned hashes unchanged. Decode CSS escapes before resolving declarations, or reject unsupported style syntax, so required rights and provenance text cannot exist only in visually hidden content.

AGENTS.md reference: AGENTS.md:L98-L105

Useful? React with 👍 / 👎.

Comment on lines +703 to +706
if tag == "a" and not hidden:
for key, value in attrs:
if key.lower() == "href" and value:
self.hrefs.append(value)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exclude anchors inside inert containers

Wrapping a registered URL in <span inert><a href="…">the same URL</a></span> still passes _validate_registry_corpus() because the visible text and destination hashes remain unchanged and this callback records every non-hidden anchor. Browsers make descendants of an inert container non-interactive, however, so readers cannot follow the purported provenance link; propagate inert ancestor state or reject inert containers around governed source anchors.

AGENTS.md reference: AGENTS.md:L98-L102

Useful? React with 👍 / 👎.

Comment on lines +1741 to +1744
for link in links:
parts.append(text[cursor:link.start])
parts.append(link.label)
cursor = link.end

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not count loaded-image alt text as visually rendered

Replacing a pinned rights sentence with ![the same sentence](a-loadable-image) still passes full registry validation because this reducer substitutes every image's alt label into the visible-text hash. When the image loads, browsers display the image rather than that alt text, so a source's provenance or redistribution boundary can disappear from reader-visible documentation without changing any integrity fixture; reject images in governed clauses or exclude their alt text from visual-content validation.

AGENTS.md reference: AGENTS.md:L98-L105

Useful? React with 👍 / 👎.

Comment on lines +1464 to +1465
visible = visible.replace("**", "").replace("__", "")
visible = visible.replace("*", "").replace("_", "")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve literal intraword underscores in rendered text

Inserting literal intraword underscores throughout a pinned sentence—for example changing The article to T_h_e_ a_r_t_i_c_l_e_—still passes _validate_registry_corpus() because this reducer deletes every underscore rather than only parsed emphasis delimiters. CommonMark renders those intraword underscores literally, so an entire rights or provenance boundary can be visibly corrupted or made unreadable while all field and complete-entry hashes continue to match; parse emphasis syntax or retain underscores that render as text.

AGENTS.md reference: AGENTS.md:L98-L105

Useful? React with 👍 / 👎.

Comment on lines +1462 to +1463
visible = _visible_html_text(visible)
visible = html.unescape(visible)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Decode character references only once

Beyond rejecting semicolonless references, replacing the start of a pinned clause with &amp;#84;he still passes full validation: a browser decodes this once and visibly displays &#84;he, but _visible_html_text() has already decoded &amp; before this second html.unescape() turns the remaining numeric-looking text into T. This lets visibly corrupted governance wording hash as the canonical sentence, so remove the second decoding pass or otherwise model the single browser-decoding stage.

AGENTS.md reference: AGENTS.md:L98-L105

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant