Data Governance & Classification
Goal
Part 4 of the AI Governance Guidelines: what data may and may not be entered into AI tools.
AI tools are only as safe as the data put into them. This part classifies institutional data into six sensitivity tiers and states, for each, what may and may not be entered into AI tools. It complements the AI Use Tiers in AI Use Tiers (which govern how much AI may be used) by governing what data AI may touch.
The governing rule: the more sensitive the data, the more it must be confined to a licensed, institution-controlled AI environment — one covered by a contract that designates the vendor a FERPA "school official," keeps institutional data as institutional property, encrypts data in transit and at rest, and prohibits training on that data. Consumer / unlicensed AI tools (free public chatbots without an institutional agreement) are appropriate only for the lowest tiers.
Throughout: "Approved AI environment" = an institution-licensed, contracted, FERPA-covered tool (e.g., BoodleBox or your institution's equivalent). "Consumer/unlicensed AI" = any public AI tool the institution has no agreement with.
4.1 Data classification tiers at a glance
| Tier | Classification | Example data types | Use with AI |
|---|---|---|---|
| D1 | Public | Already-published material: course catalogs, public web content, press releases, published syllabi, open educational resources, marketing. | Allowed in any approved AI tool. Acceptable in consumer/unlicensed AI as well. |
| D2 | Internal / Operational | Non-public, low-sensitivity: internal memos, draft course materials, meeting notes without personal data, de-identified aggregate statistics, anonymized examples. | Allowed in an approved AI environment. Avoid consumer/unlicensed tools unless data controls (no-training, account) are confirmed. |
| D3 | Proprietary/Protected | Institution-owned or third-party proprietary information and intellectual property: unpublished or pre-publication research, licensed or copyrighted content, vendor materials under NDA, internal strategy/financial/operational documents, and trade-secret-like information. | Only in a licensed, institution-controlled AI environment that does not train on your data. Must never be entered into consumer/unlicensed AI. Honor any third-party license or NDA terms first. |
| D4 | Confidential (FERPA / personal) | Student education records, grades, identifiable student work, rosters, advising notes, staff/HR personal data, licensed content (i.e. textbooks). | Only in licensed, FERPA-covered, non-training AI tools that are in an institution-controlled AI environment. Never paste into consumer/unlicensed AI. Anonymize where feasible. |
| D5 | Restricted / Regulated | Data under additional law/regulation: personally identifiable information (PII), health/HIPAA records, SSNs, financial/payment and banking data, immigration/visa, disability accommodation records, protected research data. | Default: do NOT use AI. Permitted only in a specifically approved, contractually covered environment with explicit authorization from the data owner. |
| D6 | Prohibited | Data that must never enter any AI tool: passwords/credentials/API keys, data under legal hold or protective order, export-controlled data, third parties' confidential data, anything legally barred from disclosure. | Never. No AI tool. Licensed or otherwise. No exceptions without General Counsel sign-off. |
Color cue: green = generally safe, amber = controlled environment only, red = restricted/never. When data could fall in two tiers, treat it as the higher (more sensitive) tier.
4.2 Tier definitions and rules
D1 — Public
Information already cleared for public release. There is no confidentiality risk, so it may be used freely with any approved AI tool, and is the only tier broadly appropriate for consumer/unlicensed AI. Still apply the accuracy and IP checks from AI Policy for Students.
D2 — Internal / Operational
Day-to-day institutional material that is not public but carries low risk if exposed, provided it contains no personal data. Keep it inside an approved AI environment. De-identified or aggregate data (no direct or indirect identifiers) generally belongs here rather than D3, but verify that re-identification is not possible before downgrading.
D3 — Proprietary/Protected
Information the institution or a third party owns or is contractually obligated to protect, but which is not personal data covered by FERPA. This sits a step below Confidential (D4): there is no FERPA exposure and none of the associated penalties, so it doesn't demand a FERPA-covered "school official" agreement — but it is clearly more sensitive than ordinary internal material (D2), because disclosure could harm the institution's competitive position, breach a license or NDA, or compromise others' intellectual property. The rule is therefore narrower than D2 and looser than D4: it may be used with AI, but only inside a licensed, institution-controlled environment that keeps the data as institutional property and does not train on it — never a consumer or unlicensed tool. Before sharing third-party proprietary content, confirm that the applicable license or NDA permits it.
D4 — Confidential (FERPA / personal)
Education records and personal data protected under FERPA and privacy law. This is the most common high-stakes category in teaching and advising. It may be used with AI only in a licensed environment where the vendor is contractually a "school official" under the institution's direct control, accesses only what is shared, does not train on the data, and encrypts it. It must never be entered into a consumer/unlicensed tool. When the task allows, anonymize first — metadata stripped of all direct and indirect identifiers is not PII under FERPA — and give students the ability to opt out of having their work processed by a third-party tool.
D5 — Restricted / Regulated
Data subject to additional legal regimes (HIPAA, GLBA/financial, immigration, ADA accommodation, controlled research). The default is no AI use. Where a genuine need exists, it is permitted only in an environment specifically approved for that data type, under the applicable contract/BAA, and with explicit authorization from the data owner or compliance office.
D6 — Prohibited
Data that must never be entered into any AI system under any circumstances — credentials and secrets, material under legal hold or protective order, export-controlled information, and confidential data belonging to third parties. There is no approved-tool exception; only General Counsel may authorize a narrow exception in writing.
4.3 Practical safeguards
- Default to the approved environment. For anything above D2, use only the institution's licensed, FERPA-covered AI environment — not free public chatbots.
- Anonymize before you share. De-identify student work and records whenever the task does not require identity.
- Offer opt-out. Where student work is processed by an AI tool, give students a way to decline, especially with any non-licensed tool.
- Minimize. Share only the data the task actually needs — never paste an entire roster or record when a snippet will do.
- Check the tier when in doubt. If data spans two tiers, apply the higher one. If unsure whether a tool is approved for a tier, ask IT/CIO before entering data.
- Maintain an approved-tools list. Publish which AI tools are cleared for which data tiers, with their license and FERPA status (see Governance, Roles & Review).
See also
AI Governance Guidelines · AI Use Tiers · Common Syllabus Components · Understanding data retention policies.
Questions?
Governance and security documentation is on the Trust Center: https://trust.boodlebox.ai/. Email compliance@boodle.ai for specifics, or success@boodle.ai — typical response within 8 business hours.
