Skip to content
This repository has been archived by the owner on Sep 29, 2023. It is now read-only.

Documenting data splitting procedures and best practices #14

Open
spfohl opened this issue May 14, 2021 · 0 comments
Open

Documenting data splitting procedures and best practices #14

spfohl opened this issue May 14, 2021 · 0 comments

Comments

@spfohl
Copy link

spfohl commented May 14, 2021

It would be helpful to provide documentation into how CLMBR performs data splitting and how that relates to the parameters train_end_date and val_end_date and banned_patient_file in clmbr_create_info. There should further be a discussion of best practices, and a discussion of any trade-offs, for selecting the clmbr_create_info parameters for different downstream study designs. There would ideally be examples of how to select these parameters for time splitting and patient splitting designs for different assumptions on allowed (date/time and patient) overlap between various partitions within and across pretraining and cohort-relevant partitions of the data.

Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.
Labels
None yet
Projects
None yet
Development

No branches or pull requests

1 participant