Genomics session portal 04
Kindly mark your attendance here : Click here
Kindly submit your Doubts here (Regarding any task ): Click here
Hi Interns ,
Welcome to the Internship Task Portal .
Here you will get all the intimations regarding your tasks , sessions , assignments and any updated briefing .
(The information and content provided in this portal is purely confidential and is under protected surveillance by the technical support team.)
Learn how to inspect raw sequencing reads, identify common quality issues, and prepare sequencing data for downstream genomic analysis.
- Obtain a small publicly available FASTQ dataset.
- Run quality assessment using FastQC.
- Generate a MultiQC summary report.
- Inspect per-base quality scores.
- Check adapter contamination and sequence duplication.
- Identify low-quality reads or bases.
- Perform adapter and quality trimming using a suitable tool.
- Run FastQC again after preprocessing.
- Compare the before-and-after quality reports.
- Raw FASTQ quality report.
- MultiQC report.
- Trimming/preprocessing report.
- Post-processing quality report.
- Before-and-after comparison.
- Commands or workflow documentation.
Understand the role of a reference genome and perform short-read alignment using a standard genomic alignment workflow.
- Download a suitable reference genome from a public repository.
- Inspect the FASTA reference sequence.
- Build an alignment index.
- Align paired-end sequencing reads to the reference.
- Generate a SAM/BAM alignment file.
- Convert SAM to BAM.
- Sort the BAM file.
- Index the sorted BAM file.
- Calculate basic alignment statistics.
- Reference genome information.
- Alignment command/workflow.
- Sorted BAM file or representative output.
- BAM index evidence.
- Alignment statistics.
- Short interpretation of mapping results.
Learn the basic workflow for identifying candidate single-nucleotide variants and small insertions/deletions from aligned sequencing reads.
- Use the sorted and indexed BAM file from Task 2.
- Inspect alignment depth across the reference.
- Generate candidate variant calls using a suitable variant caller.
- Generate a VCF file.
- Inspect the VCF structure.
- Review variant quality and depth fields.
- Apply basic quality filtering.
- Compare the number of variants before and after filtering.
- Raw VCF.
- Filtered VCF.
- Variant statistics.
- Depth analysis.
- Filtering criteria.
- Short explanation of the filtering process.
Learn how genomic variants can be annotated with gene and functional information and how annotation results can be organized for interpretation.
- Select a suitable public reference annotation dataset.
- Prepare the VCF for annotation.
- Annotate variants using a suitable tool or public annotation resource.
- Identify chromosome, position, reference and alternate alleles.
- Inspect gene and transcript information.
- Identify coding and non-coding variants where supported.
- Summarize the functional categories represented in the dataset.
- Create a concise annotated variant table.
- Annotated variant file.
- Variant summary table.
- Annotation-tool configuration.
- Example annotated variants.
- Brief biological interpretation.
- Reference/database information used.
Understand the basic principles of genome assembly and evaluate assembly quality using common assembly statistics.
- Use a small public sequencing dataset suitable for educational assembly.
- Perform a basic genome assembly using an appropriate assembler.
- Inspect assembled contigs.
- Calculate total assembly length.
- Calculate contig count.
- Evaluate N50 and related statistics.
- Inspect the length distribution of contigs.
- Discuss limitations of the assembly.
- Assembly FASTA file.
- Assembly statistics.
- N50 result.
- Contig count and total length.
- QUAST report or equivalent evaluation.
- Short assembly-quality interpretation.
Compare homologous genomic or protein sequences and understand how sequence alignment can reveal conserved and variable regions.
- Select 4–6 related public sequences.
- Perform multiple sequence alignment.
- Identify conserved positions.
- Identify variable regions.
- Calculate or inspect sequence identity.
- Visualize the alignment.
- Prepare a simple phylogenetic analysis if appropriate.
- Explain the biological significance of conserved regions.
- Input sequence FASTA file.
- Multiple sequence alignment.
- Alignment visualization screenshot.
- Conserved-region analysis.
- Phylogenetic tree if performed.
- Short comparative genomics report.
Understand the computational workflow used to process RNA sequencing data and identify differences in gene expression between experimental groups.
- Obtain a small public RNA-seq dataset.
- Perform raw-read quality assessment.
- Perform adapter/quality trimming where required.
- Align reads to an appropriate reference or use a suitable pseudo-alignment method.
- Generate a gene-level count matrix.
- Inspect sample-level expression patterns.
- Perform exploratory analysis such as PCA where appropriate.
- Identify candidate differentially expressed genes using a standard workflow.
- Visualize results using a volcano plot or heatmap.
- QC report.
- Processed expression/count matrix.
- Analysis workflow.
- PCA or sample clustering plot.
- Differential-expression result table.
- Volcano plot or heatmap.
- Short interpretation report.
Build a reproducible genomics workflow in which analysis steps, parameters, software versions, input data, and outputs are clearly documented.
- Create a structured project directory.
- Separate raw data, processed data, scripts and results.
- Write shell or Python scripts for repeated analysis steps.
- Record software versions.
- Create a configuration or parameter file.
- Use Git for version control.
- Create a README explaining how to reproduce the analysis.
- Generate a final workflow diagram.
- Explain how reproducibility can reduce analysis errors.
- Complete project directory.
- Analysis scripts.
- README.md.
- Software/version information.
- Git repository.
- Workflow diagram.
- Reproducibility report.
Integrate the skills learned throughout Task Portal 05 by completing a small end-to-end genomics analysis using publicly available or simulated data.
- Select an appropriate public genomics dataset.
- Document the dataset source and biological context.
- Perform initial quality control.
- Preprocess the sequencing reads where required.
- Align or map the reads to an appropriate reference.
- Generate relevant genomic outputs such as variants or expression measurements.
- Perform appropriate filtering and quality assessment.
- Annotate or characterize the resulting genomic features.
- Create at least two meaningful visualizations.
- Interpret the major computational findings.
- Document limitations and possible sources of error.
- Prepare a complete final technical report.
- Variant analysis using public sequencing data.
- Comparative analysis of related genomes.
- Small-scale bacterial genome analysis.
- RNA-seq expression comparison using public data.
- Sequence conservation analysis.
- Genome assembly and quality assessment.
- Publicly available population-genomics dataset exploration.
- Project Title
- Research Question / Objective
- Dataset Description
- Reference Genome / Annotation
- Tools and Software Versions
- Computational Methodology
- Quality Control Results
- Analysis Results
- Visualizations
- Biological Interpretation
- Limitations
- Conclusion
- References
- GitHub Repository Link
- Complete source/data-processing workflow.
- Scripts and configuration files.
- Relevant processed datasets.
- Analysis results.
- Figures and visualizations.
- GitHub repository.
- Complete final report in PDF.
- Project presentation/demo screenshots.
After completing all 9 tasks, organize your complete work into one main folder. Each task should have its own subfolder containing the relevant files and evidence.
Use the following naming format:
- Source code and scripts
- FASTQ/FASTA files where appropriate and legally shareable
- Processed data/results
- VCF/BAM/other analysis outputs where appropriate
- QC reports
- Graphs and visualizations
- Terminal/workflow screenshots
- Configuration files
- README documentation
- GitHub repository links
- Task-specific reports
Add one final PDF report to the main folder covering your overall experience and technical work across Task Portal 05.
- Intern name
- Internship/program details
- Tasks completed
- Tools and technologies used
- Project summaries
- Methodology
- Results and visualizations
- Challenges and solutions
- Key learning outcomes
- GitHub/project links
- Final mini-project documentation
Upload the complete submission folder to Google Drive or another suitable cloud-storage platform.
Share the complete folder with the following official email address:
Make sure the folder permissions allow the internship evaluation team to access the submitted files.
After sharing the folder, copy the Google Drive share link and submit the link through the designated internship task submission form/portal.
- ☐ All 9 task folders created
- ☐ Required files added to each task folder
- ☐ Scripts and source code included
- ☐ Results and reports included
- ☐ Screenshots/visualizations included
- ☐ GitHub repository link included
- ☐ Final PDF report included
- ☐ Complete folder uploaded to Google Drive
- ☐ Folder shared with support@thenexoragroup.com
- ☐ Access permission checked
- ☐ Shareable folder link copied
- ☐ Folder link submitted through the task portal
Do not upload patient-identifiable genomic information, private research datasets, passwords, API keys, access tokens, private credentials, or confidential institutional data. Use public or simulated datasets unless you have explicit authorization to share other data.
Please submit one properly organized folder containing your complete Task Portal 05 work. Avoid sending individual files separately unless specifically requested by the internship team.
