Why the metadata answers the questions and the files do not
Clients arrive wanting the documents and leave wanting the metadata, because the metadata is what answers their questions and the documents come with problems nobody wants.
The questions are these. Which courses generate the most study material, and at which institutions. How curriculum coverage differs between universities teaching the same subject. Which textbooks and topics dominate a field. Where demand for tutoring or supplementary content concentrates. Every one of those is answerable from institution, course and volume data without touching a single uploaded file.
For an education company planning where to build content, that map is the brief. For a publisher it shows which courses actually use which material. For a university it shows what its own students are circulating, which is occasionally uncomfortable and always informative.
The files themselves answer almost nothing additional and carry real exposure: institutional copyright on the assignments, student authorship on the notes, and academic integrity questions on the solutions. A dataset of them is not something a company can defend, and we would rather say that at the scoping call than build it and watch somebody else find out.
The last consideration is the platform's own position, which is worth reading in its crawl rules before deciding what is reasonable. They are specific about what is open and what is not, and staying inside that line is also what keeps access stable.