IMPORTANT 
LAB WILL BE CLOSED ON OCT 8

The Nucleomics Core Facility will be closed on Thursday October 8th to celebrate the 30Y anniversary of the VIB.
Please take this into account before sending/bringing your samples.

FAQ

Sample preparation

Does VIB Nucleomics core offer Sanger sequencing ?

We do not do Sanger sequencing here, only Next-and Third Generation Sequencing. For Sanger sequencing, please feel free to contact our colleagues at NSF Antwerp: https://cmn.sites.vib.be/en/neuromics-support-facility#/

Do you provide DNA or RNA extraction as a service?

Currently we do not provide DNA or RNA extraction as a service to our customers. However, we can be an intermediate for outsourcing your samples for RNA/DNA extraction to a collaborative service facility with which we have good experiences. Please contact us for more information.

Can you recommend a RNA/DNA extraction protocol?

Currently, we do not impose a specific protocol for RNA/DNA extractions. Our advice would be to use a protocol which works best in your hands. Please keep in mind that your RNA/DNA samples must meet specific quality criteria in order for us to be able to guarantee good quality data. Check the Sample Submission Page for more details.

 

What are the requirements for my DNA sample?

You can find more details about the requirements on our Sample Submission Page.

What are the requirements for my RNA sample?

You can find more details about the requirements on our Sample Submission Page.

What if I have only a very low amount of RNA?

The standard protocols for the majority of the assays that we offer start with about 100-200 ng total RNA. If you have less, there are ways of doing more extensive amplifications. For this we use the NuGEN Ovation Pico WTA system. Although the specifications of this kit say that the minimal amount is 500 pg, we have tested the reproducibility and linearity between different amounts of starting material. Our advice is not to go below 5-10 ng. However, if you do not have enough sample material, and still want to continue, you often have no other choice than to accept a bias. We could do a QC for your samples on an Agilent Bioanalyzer Pico chip, which is able to also quantitatively measure your sample with a lower limit of 50 pg/µL. If you would like to discuss this in more detail, including advice on platform type, please contact nucleomics@vib.be to make an appointment for a meeting at our facility.

How should I bring in my DNA or RNA samples?

You can bring your samples in any working day between 9-12h and 13-17h. If you ship your samples by courier, please ship them at the beginning of the week on plenty of dry-ice. Before you bring/ship the samples, please make sure that you send us a filled out DNA or RNA Sample Submission Form. We will check the quality of the samples and contact you again before we start the actual assays.

Check the Sample Submission Page for more details.

Bioinformatics

Do you still have the FASTQ files from my run ?

We keep sequencing data for at least 3 months on our servers. 

If your sequencing data are not yet deleted, the recovery of data >3months from our archives comes at a flat administrative fee of 75€ for data from an average project.

Why do I have multiple files per sample ? and how to handle it ?

On Illumina instrument, each sample can be distributed across multiple lanes (4 on NextSeq500, 2 on the NextSeq2000 and AVITI24, and up to 4 on the NovaSeq6000). The lane can be identified in the filename as L001, L002, L003, L004.

If the lanes could be considered as technical replicates for the sequencing, they should not be handled as biological replicates of your experiment.

We used to map per lane (optimal for parallelization of the pipeline), then we merge the bam files using samtools command as follow:

samtools merge [options] sample_S1_L001_R1.merged.bam sample_S1_L001_R1.bam sample_S1_L002_R1.bam ...

Another approach would be to merge the fastq files from different lanes per sample before starting the pipeline. You can also sum up the counts from different lanes per sample after the counting.

Note that you get a double number of files per sample if the sequencing is paired-end. The fastq files are splitted in Read 1 (forward, "_R1" in the filename) and the second is Read 2 (Reverse, "_R2" in the filename). You only get Read1 files for single-end sequencing.

Why my RNA-seq samples where sequenced twice ?

For RNA-seq projects, we often sequence twice in order to correct for the pooling. 

This resequencing approach allows us to be as close as possible to the equimolar expectations when you combine the output of all runs. 

If the output per lane and per run could be considered as technical replicates for the sequencing, they should not be handled as biological replicates of your experiment. 
Thus, for your data processing, you need to merge the data for each sample of both runs and both lanes (see methods in the previous question).

I don’t find my data using Basespace CLI?

We recently moved our Basespace data from US to EU host server. You will need to specify it in your command.

bs auth –api-server=https://api.euc1.sh.basespace.illumina.com/

Can I attend a data analysis session?

It is not our usual practice that customers follow us when we do the analysis. For most of the analysis, we use freeware tools such as R/Bioconductor. You will also get a full analysis report describing the complete analysis and which tools and packages that we have used. With this, you will have sufficient information on how we do these kinds of analyses. However, since we use R, some programming skills are required. In case you really look for additional bioinformatics training, we can recommend the courses that are organized by VIB Bioinformatics Training.

How can I motivate the use of ‘uncorrected’, but very stringent, p-values?

Once p-values are computed, we define a p-value-based criterion for selecting genes. We need to make a trade-off between precision and recall, or otherwise false discovery rate and statistical power. Simply using a cut-off of 0.05 would result in unacceptably many false positives, given the large number of statistical tests performed. Multiple testing criteria take in account the distribution of p-values and are developed to limit the number of false positives, however at the expense of a higher number of false negatives. We evaluated several multiple testing strategies (Benjamini and Hochberg 1995; Storey 2003; Scheid and Spang 2005), some of them combined with precursory filtering based on overall variance (Bourgon et al. 2010), or starting from p-values computed from other statistics than the (moderated) t-statistic (Hong et al. 2006). For some experiments, all approaches returned at most a few genes, even when mild cut-offs were used. It is a well-recognized issue that adjustment for multiple testing can result in low statistical power for micro-array studies (Bourgon et al. 2010). An evident reason is the low number of samples combined with a high number of tests, but also other aspects such as gene correlation can play an important role (Efron 2007). Because the selected genes are usually independently validated afterwards, we choose for a selection criterion that returns less false positives than the simple 0.05-cut-off and less false negatives than the general multiple testing approaches. Therefore we adopt the criterion that was used during the elaborate MAQC-I study and select genes based on p-value < 0.001, further constrained to genes with an absolute fold-change > 2 (MAQC Consortium 2006).

References

  • Benjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological), Vol. 57, No. 1, pp. 289-300. (1995)
  • Bourgon R, Gentleman R, Huber W. Independent filtering increases detection power for high-throughput experiments. Proceedings of the National Academy of Sciences of the United States of America, Vol. 107, No. 21. (May), pp. 9546-9551. (2010)
  • Efron B. Size, power and false discovery rates. The Annals of Statistics, Vol. 35, No. 4. (August), pp. 1351-1377. (2007)
  • Hong F, Breitling R, McEntee C.W, Wittner B.S., Nemhauser J.L, Chory J. RankProd: a bioconductor package for detecting differentially expressed genes in meta-analysis. Bioinformatics, Vol. 22, No. 22. (November), pp. 2825-2827. (2006)
  • MAQC Consortium. The MicroArray Quality Control (MAQC) project shows inter- and intraplatform reproducibility of gene expression measurements. Nature Biotechnology, Vol. 24, No. 9. (September), pp. 1151-1161. (2006)
  • Scheid S, Spang R. twilight; a Bioconductor package for estimating the local false discovery rate. Bioinformatics, Vol. 21, No. 12. (June), pp. 2921-2922. (2005)
  • Storey J.D. The positive false discovery rate: a Bayesian interpretation and the q -value. The Annals of Statistics, Vol. 31, No. 6. (December), pp. 2013-2035.(2003)

How can I run a functional analysis myself using DAVID?

You may try out the user-friendly DAVID tool for a quick functional analysis. 

  • Click on “Shortcut to DAVID Tools > Functional Annotation” in the menu
  • To upload the gene list, go to your Statistical Table, select your differentially expressed genes (e.g. all up-regulated genes from first contrast are with “1” in column G) and copy paste the Ensembl Gene ID column (1st column) of the filtered rows.
  • Enter your input as gene list:
    • Paste the list of genes in the box at the left entitled “A: Paste a list”
    • Select the identifier to be “ENSEMBL_GENE_ID”
    • Select the list type to be “Gene List”
    • Click on “Submit List”
  • Enter your background:
    • Clear active filters on your Statistical Table and select the Gene IDs (column 1) of all genes.
    • Go back to "Upload" tab
    • Paste the list of all genes in the box at the left entitled “A: Paste a list”
    • Select the identifier to be “ENSEMBL_GENE_ID”
    • Select the list type to be “Background”
    • Click on “Submit List”
  • Configure your search:
    • Go to "Start Analysis" > "Functional Annotation" tools
    • Choose the databases and ‘vocabulary’ in which the functional characterization is formulated. Usually we deselect “Check defaults” and select our own choices: Click on "Gene_Ontology" and select all levels ending with “GOTERM_BP_3”, “GOTERM_BP_4”, or “GOTERM_BP_5” (these are the most detailed, low level concepts). Click on "Pathways" and select “KEGG PATHWAY”, "REACTOME PATHWAY" and "WIKIPATHWAYS".
  • Finalize your search by clicking at the bottom on “Functional Annotation Clustering”.

Interpret the output: Several blocks appear that cluster together closely related functional terms/pathways. The block at the top is the most important one, indicated by the enrichment scores. To get more information, you can click on the terms. Clicking on the Kegg pathway terms (if present in the list) shows the pathways and annotates genes in your list by a blinking star. In this sense, you can get an idea about the function of genes in your list.

How can I run a functional analysis using WebGestalt ?

For the use of WebGestalt, you should :

1- Methods of Interest : Over-Representation Analysis (ORA) [default]

2- select Homo sapiens as organism of interest 

3- Functional Databases (click on the + button to add multiple databases): 

  • Gene ontology, then select below Biological Process noRedundant
  • Pathways, then Reactome
  • Pathways, then WikiPathways

4- To upload the gene list, go to your Statistical Table, select your differentially expressed genes (e.g. all up-regulated genes from first contrast are with “1” in column G) and copy paste the Ensembl Gene ID column (1st column) of the filtered rows.

5- Gene List: use the Ensembl Gene ID as gene ID type. Note that Gene Symbol can work as well but they can be less consistent.

6- Select “Genome” as Reference Set in the Reference Gene list section if you have a total RNA-seq, or select "Genome protein-coding" if you have a mRNA-seq experiment.

7- Type Submit, the results will be loaded in a separate tab on your web browser.

8- Wait for the results…. (usually less than a minute)

9- When you are satisfied with the results, you can click on “Result Download” link to save the results in a compressed archive (zip) to save it on your computer.

How can I run a GSEA analysis using WebGestalt ?

For the use of GSEA of WebGestalt, you should  :

1- select the Methods of Interest : Gene Set Expression Analysis

2- select Homo sapiens as organism of interest 

3- Functional Databases : You can only select one database at a time. For example you can choose Gene ontology, then select below Biological Process noRedundant

4- To upload the ID list, you will need to upload a ranked list file with two columns: gene IDs and the scores separated by tab (or paste IDs and the scores separated by tab in the text box). Maximum upload file size is 5MB.

To generate this ranking file, go to your Statistical Table:

  • add an empty column to your table,
  • identify the columns for logFC and statistical values (PValue or FDR) for the contrast of interest,
  • let's say we have column C for logFC and column E for FDR values, we can compute the score as the signed log of the FDR using this formula 

=SIGN(C)*-LOG(E)

  • hide the columns between the gene ID (column A) and the Score column, and copy paste the values of the two columns in one empty text file. Do select all the genes (no need to take the header).
  • update the text file to change ',' in '.' for the decimals
  • save the file with .rnk and upload this file to the tool.

5- Ensembl Gene ID should be identified as gene ID type. 

7- Type Submit, the results will be loaded in a separate tab on your web browser.

8- Wait for the results…. (usually less than a minute)

9- When you are satisfied with the results, you can click on “Result Download” link to save the results in a compressed archive (zip) to save it on your computer.