2026-05-21 Better Sample Data Meeting notes

2026-05-21 Better Sample Data Meeting notes

 Date

Feb 24, 2026

 Participants

@Autumn Faulkner, @Shelley Doljack , @Tod Olson

Regrets: @Charlotte Whitt @Lee Braginsky @Yogesh Kumar (Deactivated)

 Goals

  • Brainstorm next steps

 Discussion topics

Time

Item

Notes

 

Time

Item

Notes

 

 

Anonymization scripts status

  • Context: Would like to complete work on anonymization scripts, run them on an MSU dataset as a means of demonstrating successful PPI removal for MSU’s Risk Management office

  • Lee needs volunteer developer time to complete the scripts

  • Update project call to include use of AI as a possible tool

  • Will post call on Slack channels

 

 

Generating a dataset instead of seeking a contributed/anonymized one

  • Given the difficulty we’ve had making headway on a contributed dataset, is generating random data the better approach?

  • For users and bib data, generating random user profiles would not be too challenging

  • It’s historical data that is the problem, and the record of cross-app interactions

    • I.e., loan rules and check-outs related to instances

    • Fiscal year budgets and ledgers associated with certain orders

  • Could AI be a tool for generating the more complex data and data interactions, if fed the appropriate JSON schema info?

    • Shelley would like to investigate, as time permits

 

 

Presentation from GBV/Index Data colleagues in June

  • A look at their approaches to populating reference environments with good sample data

  • MODINVUP-184: GBV. MIU. Add sample data to be populated in the reference environmentsClosed

 

 Action items