Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Addition of 2025_GnecchiRuscone_CarpathianBasinHunPeriod #243

Closed
wants to merge 0 commits into from

Conversation

ltcrod
Copy link
Contributor

@ltcrod ltcrod commented Jan 29, 2025

PR Checklist for a new package submission

  • The package does not exist already in the community archive, also not with a different name.
  • The package title in the POSEIDON.yml conforms to the general title structure suggested here: <Year>_<Last name of first author>_<Region, time period or special feature of the paper>, e.g. 2021_Zegarac_SoutheasternEurope, 2021_SeguinOrlando_BellBeaker or 2021_Kivisild_MedievalEstonia.
  • The package is stored in a directory that is named like the package title.

  • The package is complete and features the following elements:
    • Genotype data in binary PLINK format (not EIGENSTRAT format).
    • A POSEIDON.yml file with not just the file-referencing fields, but also the following meta-information fields present and filled: poseidonVersion, title, description, contributor, packageVersion, lastModified (see here for their definition)
    • A reasonably filled .janno file (for a list of available fields look here and here for more detailed documentation about them).
    • A .bib file with the necessary literature references for each sample in the .janno file.
  • Every file in the submission is correctly referenced in the POSEIDON.yml file and there are no additional, supplementary files in the submission that are not documented there.
  • Genotype data, .janno and .bib file are all named after the package title and only differ in the file extension.
  • The package version in the POSEIDON.yml file is 1.0.0.
  • The poseidonVersion of the package in the POSEIDON.yml file is set to the latest version of the Poseidon schema.
  • The POSEIDON.yml file contains the corresponding checksums for the fields genoFile, snpFile, indFile, jannoFile and bibFile.
  • There is either no CHANGELOG file or one with a single entry for version 1.0.0.

  • The Publication column in the .janno file is filled and the respective .bib file has complete entries for the listed mentioned keys.
  • The .janno file does not include any empty columns or columns only filled with n/a.
  • The order of columns in the .janno file adheres to the standard order as defined in the Poseidon schema here.
  • The .janno and the .ssf files are not fully quoted, so they only use single- or double quotes ("...", '...') to enclose text fields where it is strictly necessary (i.e. their entry includes a TAB).

  • The package passes a validation with trident validate --fullGeno.

  • Large genotype data files are properly tracked with Git LFS and not directly pushed to the repository. For an instruction on how to set up Git LFS please look here. If you accidentally pushed the files the wrong way you can fix it with git lfs migrate import --no-rewrite path/to/file.bed (see here).

@stschiff
Copy link
Member

Thanks @ltcrod could you briefly clarify/edit the following:

  1. Why does this PR contain also edits to the YAML file of a previous package (2024_GnecchiRuscone_CarpathianBasinAvarPedigrees)?

  2. So this is unpublished data, correct?

  3. The Alternative_ID column is strangely in most cases just the same as Poseidon_ID. Could you clarify? I think you can just remove all values that are the same as the Poseidon_ID as they don't add anything, do they?

  4. I think for the Data_Preparation_Pipeline_URL you can add a link to the department-pipeline Eager config file, if that was used.

  5. What about fields Endogenous, Coverage_on_Target_SNPs, could you fill those too, by any chance?

  6. I would suggest that empty columns can be removed, as indicated also in the checklist above. You can also now use the new --jannoRemoveEmpty option for trident rectify for that.

@ltcrod
Copy link
Contributor Author

ltcrod commented Jan 31, 2025

Hi @stschiff thanks for the review.

  1. Misclic when updating locally, will open a new PR without it if necessary
  2. Yes, unpublished
  3. yes, it can be left blank for all but two then
  4. Yes
  5. I have always left the Endogenous blank for samples with no WG screening, can add something. Coverage_on_Target_SNPs can definitely be filled
  6. Thanks for the suggestion, that can be done

@stschiff
Copy link
Member

stschiff commented Jan 31, 2025

For 1, please don't make a new PR. Just commit a manual undo, and it will be fine, I think. You just want to check on the "Files changed" tab in this PR that only files within this package show up.

@stschiff
Copy link
Member

Any update here, @ltcrod ?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

Successfully merging this pull request may close these issues.

2 participants