Why the lifecycle is the part worth building
Almost every brief that arrives here can be split into two halves, and the halves need completely different answers.
The first half - find documents matching these criteria, give me their metadata and text - is already solved by the published interface. Quoting a scraping project for it would be charging for a wrapper around something free, and we do not do that.
The second half is the lifecycle, and it is genuinely unsolved. A rule that was proposed in one year, drew comments, was finalised in another and amended after that is one regulatory story told across several documents. Assembling that chain, with dates, stages and the relationships between documents, is what turns a document archive into something a compliance or policy team can act on.
The third piece is cross-source joining. Rulemaking connects to the comment dockets, to the codified regulations and to agency guidance published elsewhere, and the value in a regulatory dataset usually sits in those joins rather than in any single source.
The fourth is timing. Effective dates, comment deadlines and publication dates are three different dates, and a dataset that carries one of them and calls it the date will eventually cause somebody to miss a deadline.