Genomics session portal 06

Kindly mark your attendance here : Click here
Kindly submit your Doubts here (Regarding any task ): Click here
Hi Interns ,
Welcome to the Internship Task Portal .
Here you will get all the intimations regarding your tasks , sessions , assignments and any updated briefing .
(The information and content provided in this portal is purely confidential and is under protected surveillance by the technical support team.)
Genomics Internship – Task Portal 05
☑ TASK 1: Genome Annotation & Gene Prediction INTERMEDIATE
🎯 Objective

Learn how structural genome annotation identifies genes, coding regions and other genomic features in an assembled sequence.

🧪 Practical Work
  • Select a small publicly available bacterial or other suitable genome assembly.
  • Inspect the FASTA sequence and basic assembly statistics.
  • Run a suitable genome annotation tool such as Prokka or an equivalent workflow.
  • Identify predicted genes, CDS features, rRNAs and tRNAs where supported.
  • Inspect the generated GFF/GBK and protein FASTA files.
  • Calculate the number of annotated features.
  • Summarize the functional annotation categories.
  • Create a concise genome annotation report.
💻 Commands / Workflow to Practice
prokka --outdir annotation_output \ --prefix sample_annotation \ assembly.fasta # Inspect annotation files ls annotation_output/ # Count CDS features grep -c "CDS" annotation_output/sample_annotation.gff
📌 Submission Evidence
  • Input genome FASTA.
  • Annotation output files.
  • Feature-count summary.
  • GFF/GenBank inspection screenshots.
  • Functional annotation summary.
  • Short technical report.
☑ TASK 2: Promoter, Motif & Regulatory Sequence Analysis INTERMEDIATE
🎯 Objective

Explore DNA sequence motifs and regulatory regions to understand how sequence patterns can be associated with gene regulation.

🧪 Practical Work
  • Select a suitable gene or small set of genes from a public organism.
  • Retrieve upstream promoter-region sequences.
  • Perform sequence composition analysis.
  • Search for known regulatory motifs using a suitable resource.
  • Identify recurring motifs and their positions.
  • Compare motif occurrence between selected sequences.
  • Visualize motif locations.
  • Discuss limitations of motif-based interpretation.
💻 Commands / Workflow to Practice
# Example sequence inspection seqkit stats promoter_sequences.fasta # Search a simple motif grep -n "TATA" promoter_sequences.fasta # Example workflow: # 1. Retrieve promoter sequences # 2. Run a motif-search tool/database # 3. Export motif locations # 4. Visualize and summarize the results.
Use publicly available sequences and clearly document the organism, gene identifiers and sequence source.
📌 Submission Evidence
  • Promoter FASTA file.
  • Motif-search output.
  • Motif-position table.
  • Sequence visualization.
  • Interpretation and limitations.
☑ TASK 3: Metagenomics Read Classification & Microbial Profiling INTERMEDIATE
🎯 Objective

Understand how sequencing reads from mixed microbial communities can be classified and summarized into a taxonomic profile.

🧪 Practical Work
  • Obtain a small public metagenomic dataset.
  • Perform basic FASTQ quality assessment.
  • Remove adapters or low-quality reads if required.
  • Classify reads using a suitable educational metagenomics workflow.
  • Generate a taxonomic abundance table.
  • Identify dominant taxonomic groups.
  • Visualize the community composition.
  • Discuss classification limitations and database dependence.
💻 Example Workflow
fastqc metagenome_R1.fastq.gz fastqc metagenome_R2.fastq.gz multiqc . # Example classification workflow kraken2 --db YOUR_DB \ --paired metagenome_R1.fastq.gz metagenome_R2.fastq.gz \ --report kraken_report.txt \ --output kraken_output.txt # Convert/summarize the report using an appropriate visualization tool.
📌 Submission Evidence
  • QC report.
  • Classification report.
  • Taxonomic abundance table.
  • At least one community-composition visualization.
  • Short interpretation report.
☑ TASK 4: Functional Enrichment & Pathway Analysis INTERMEDIATE
🎯 Objective

Learn how a list of genes can be connected to biological processes, molecular functions and pathways using public annotation resources.

🧪 Practical Work
  • Obtain a public gene list or generate one from a suitable educational dataset.
  • Standardize gene identifiers where required.
  • Select an appropriate organism and annotation database.
  • Perform Gene Ontology enrichment.
  • Perform pathway enrichment using a suitable resource.
  • Rank enriched categories using an appropriate statistical measure.
  • Create a bar plot or dot plot of important categories.
  • Explain how enrichment differs from simple gene-list inspection.
💻 Example R Workflow
library(clusterProfiler) library(org.Hs.eg.db) ego <- enrichGO( gene = gene_ids, OrgDb = org.Hs.eg.db, keyType = "ENTREZID", ont = "BP", pAdjustMethod = "BH" ) dotplot(ego)
📌 Submission Evidence
  • Input gene list.
  • Database/organism information.
  • Enrichment result table.
  • GO/pathway visualization.
  • Biological interpretation.
☑ TASK 5: Genome-Wide Association Study Data Exploration INTERMEDIATE+
🎯 Objective

Understand the basic structure of genotype-phenotype association data and explore how statistical signals are visualized in a GWAS workflow.

🧪 Practical Work
  • Use a public or simulated GWAS dataset.
  • Inspect genotype, phenotype and variant metadata.
  • Perform basic data-quality checks.
  • Review missingness and allele-frequency information.
  • Run or inspect an educational association-analysis output.
  • Apply an appropriate multiple-testing correction or significance threshold.
  • Create a Manhattan plot and QQ plot.
  • Identify candidate association signals without presenting them as clinical conclusions.
💻 Example Workflow
# Example PLINK-style commands plink --bfile study_data \ --missing \ --freq \ --hardy \ --out qc_summary # Example association analysis plink --bfile study_data \ --pheno phenotype.txt \ --assoc \ --out gwas_results # Generate plots using R/Python or a suitable GWAS visualization package.
Educational use only: Use public or simulated datasets. Do not make medical or clinical claims from the exercise.
📌 Submission Evidence
  • Dataset description.
  • Quality-control summary.
  • Association result table.
  • Manhattan plot.
  • QQ plot.
  • Short statistical interpretation.
☑ TASK 6: Epigenomics & DNA Methylation Data Analysis INTERMEDIATE+
🎯 Objective

Explore how DNA methylation data can be processed, summarized and compared between biological groups.

🧪 Practical Work
  • Obtain a public methylation dataset or simulated methylation matrix.
  • Inspect CpG-level measurements and sample metadata.
  • Perform basic data-quality checks.
  • Visualize methylation distributions.
  • Compare methylation patterns between two groups.
  • Identify candidate differentially methylated sites or regions using a suitable method.
  • Create a heatmap or clustering visualization.
  • Discuss biological and technical sources of variation.
💻 Example R Workflow
meth <- read.csv( "methylation_matrix.csv", row.names = 1 ) summary(meth) boxplot( meth, las = 2, main = "Methylation Distribution" ) # Continue with an appropriate differential # methylation analysis package/workflow.
📌 Submission Evidence
  • Methylation dataset.
  • Sample metadata.
  • Quality-control summary.
  • Statistical results.
  • Heatmap/boxplot or equivalent visualization.
  • Interpretation report.
☑ TASK 7: Single-Cell RNA-Seq Exploratory Analysis INTERMEDIATE+
🎯 Objective

Learn the major computational steps used to explore single-cell RNA sequencing data and identify cell populations.

🧪 Practical Work
  • Use a small public single-cell expression dataset or an educational matrix.
  • Inspect cells, genes and metadata.
  • Apply basic quality-control filters.
  • Normalize the expression data.
  • Identify highly variable genes.
  • Perform dimensionality reduction using PCA and/or UMAP.
  • Cluster cells.
  • Inspect marker genes for selected clusters.
  • Create a cell-cluster visualization.
💻 Example Seurat Workflow
library(Seurat) obj <- CreateSeuratObject( counts = counts_matrix ) obj <- NormalizeData(obj) obj <- FindVariableFeatures(obj) obj <- ScaleData(obj) obj <- RunPCA(obj) obj <- FindNeighbors( obj, dims = 1:10 ) obj <- FindClusters(obj) obj <- RunUMAP( obj, dims = 1:10 ) DimPlot( obj, reduction = "umap", label = TRUE )
Document dataset source, species, filtering thresholds, software versions and clustering parameters.
📌 Submission Evidence
  • Dataset and metadata.
  • QC summary.
  • PCA/UMAP visualization.
  • Cluster results.
  • Marker-gene table for selected clusters.
  • Short interpretation report.
☑ TASK 8: Protein Sequence Analysis & Functional Prediction INTERMEDIATE
🎯 Objective

Connect genomic information with protein-level analysis by examining sequence features, domains and predicted functional characteristics.

🧪 Practical Work
  • Select 2–4 related protein sequences from a public database.
  • Calculate sequence length and amino-acid composition.
  • Perform pairwise or multiple sequence alignment.
  • Search for conserved domains using a public domain resource.
  • Identify conserved residues or regions.
  • Compare predicted functional features between proteins.
  • Visualize the alignment and domain organization.
  • Write a short structure-function interpretation.
💻 Tools / Commands to Practice
seqkit stats proteins.fasta mafft --auto proteins.fasta > proteins_aligned.fasta # Example domain-analysis workflow: # 1. Submit/query protein sequences against # a suitable public domain database. # 2. Record domain names and coordinates. # 3. Compare conserved regions between proteins.
📌 Submission Evidence
  • Protein FASTA sequences.
  • Alignment file.
  • Domain-analysis output.
  • Conserved-region table.
  • Alignment/domain visualization.
  • Functional interpretation.
☑ TASK 9: Genomics Integration Project — Multi-Omics Mini Study INTERMEDIATE+
🎯 Objective

Integrate multiple genomics concepts into a small reproducible project that connects sequence, expression or functional information into one analytical story.

🧪 Practical Work
  1. Select a public or simulated dataset appropriate for an educational analysis.
  2. Define one clear biological research question.
  3. Document the dataset source and biological context.
  4. Perform appropriate quality control.
  5. Process at least two complementary data layers.
  6. Perform statistical or computational analysis appropriate to the dataset.
  7. Integrate the outputs into a combined result table.
  8. Create at least three meaningful visualizations.
  9. Identify major patterns and possible biological relationships.
  10. Discuss limitations, assumptions and potential sources of error.
  11. Prepare a reproducible workflow and final technical report.
🧬 Suggested Project Topics
  • Variant and functional annotation integration.
  • RNA-seq and pathway enrichment analysis.
  • Metagenomic profiling and functional interpretation.
  • Genome annotation and comparative analysis.
  • DNA methylation and gene-expression exploration.
  • Single-cell clustering and pathway characterization.
  • Protein conservation and genomic annotation integration.
📊 Required Final Report Sections
  1. Project Title
  2. Research Question / Objective
  3. Dataset Description
  4. Data Sources and References
  5. Tools and Software Versions
  6. Computational Methodology
  7. Quality Control
  8. Analysis Results
  9. Integrated Results
  10. Visualizations
  11. Biological Interpretation
  12. Limitations
  13. Conclusion
  14. References
  15. GitHub Repository Link
📌 Final Submission Evidence
  • Complete analysis workflow.
  • Scripts/notebooks.
  • Processed datasets where legally shareable.
  • Results and tables.
  • At least three figures.
  • GitHub repository.
  • Complete final report in PDF.
  • Project presentation/demo screenshots.
Important: Use publicly available or simulated datasets for this educational project. Do not upload patient-identifiable genomic data, private sequencing data, confidential research data, passwords, API keys or access tokens.
📤 TASK PORTAL 05 — SUBMISSION METHOD

After completing all 9 tasks, organize your complete work into one main folder. Each task should have its own subfolder containing the relevant files and evidence.

1️⃣ Create the Main Submission Folder

Use the following naming format:

Genomics_Internship_Task_Portal_05_YourName
2️⃣ Create Separate Folders for All 9 Tasks
Genomics_Internship_Task_Portal_05_YourName/ ├── Task_1_Genome_Annotation/ ├── Task_2_Regulatory_Motif_Analysis/ ├── Task_3_Metagenomics_Profiling/ ├── Task_4_Functional_Enrichment/ ├── Task_5_GWAS_Data_Analysis/ ├── Task_6_Epigenomics_Analysis/ ├── Task_7_Single_Cell_RNA_Seq/ ├── Task_8_Protein_Sequence_Analysis/ └── Task_9_Multi_Omics_Mini_Project/
3️⃣ Add All Required Evidence
  • Source code, scripts and notebooks.
  • FASTQ/FASTA files where appropriate and legally shareable.
  • Processed data and analysis results.
  • QC reports.
  • Graphs and visualizations.
  • Terminal/workflow screenshots.
  • Configuration files.
  • README documentation.
  • GitHub repository links.
  • Task-specific reports.
4️⃣ Add Your Final Internship Report

Add one final PDF report to the main folder covering your overall experience and technical work across Task Portal 05.

  • Intern name
  • Internship/program details
  • Tasks completed
  • Tools and technologies used
  • Project summaries
  • Methodology
  • Results and visualizations
  • Challenges and solutions
  • Key learning outcomes
  • GitHub/project links
  • Final multi-omics mini-project documentation
5️⃣ Upload the Complete Folder

Upload the complete submission folder to Google Drive or another suitable cloud-storage platform.

6️⃣ Share the Folder With the Internship Team

Share the complete folder with the following official email address:

📧 support@thenexoragroup.com

Make sure the folder permissions allow the internship evaluation team to access the submitted files.

7️⃣ Submit the Folder Link

After sharing the folder, copy the Google Drive share link and submit the link through the designated internship task submission form/portal.

8️⃣ Final Submission Checklist
  • ☐ All 9 task folders created
  • ☐ Required files added to each task folder
  • ☐ Scripts/notebooks included
  • ☐ Results and reports included
  • ☐ Screenshots/visualizations included
  • ☐ GitHub repository link included
  • ☐ Final PDF report included
  • ☐ Complete folder uploaded to Google Drive
  • ☐ Folder shared with support@thenexoragroup.com
  • ☐ Access permission checked
  • ☐ Shareable folder link copied
  • ☐ Folder link submitted through the task portal
🔐 DATA & SECURITY NOTICE:
Do not upload patient-identifiable genomic information, private research datasets, passwords, API keys, access tokens, private credentials, or confidential institutional data. Use public or simulated datasets unless you have explicit authorization to share other data.
📌 Important:
Please submit one properly organized folder containing your complete Task Portal 05 work. Avoid sending individual files separately unless specifically requested by the internship team.