Full Time
Php 1500
TBD
Oct 2, 2025
Project Brief (Companies Only – Next Generation Sequencing Focus)
Important Note (Read First)
For this project, longer lists with mandatory attributes are more valuable than shorter lists with many nice-to-haves.
If you must choose, always prioritize volume with mandatory fields over completeness with nice-to-haves.
Accuracy still matters — no duplicates, no fake LinkedIns, no broken domains.
Nice-to-have attributes are welcome, but they should not reduce the overall scale of the dataset.
What We Need Collected
Companies
Mandatory:
• Company domain (website URL) OR
• Company name
Nice-to-have (if available):
• Company name
• Industry / sub-industry classification
• Location (city, state, country)
• Employee size (only include companies with 50–3,000 employees, reported in ranges:
• 51–200
• 201–500
• 501–1,000
• 1,001–3,000)
•
Volume & Expectations
• NGS-focused dataset: minimum goal of 5,000 companies globally, with emphasis on:
• Primary regions: North America, Europe
• Secondary regions: Australia, Israel, Singapore
• Within the dataset, create sub-lists by research/application area (e.g., oncology research, immunotherapy, vaccine development, cell therapy, biomarker discovery).
???? If you cannot fully reach the target, you must report blockers (e.g., lack of data sources, technical limits) in a timely manner.
Data Quality
• Companies must be within 50–3,000 employees only
• No duplicates, no broken domains, no fake LinkedIn profiles
• Prioritize Company Domain whenever available
• Must clearly fit into NGS ecosystem (see keyword framework and exclusion rules below)
Target Company Indicators
Key Search Keywords
• Primary NGS: Next-generation sequencing, Genomics, Transcriptomics, Single-cell RNA-seq, 10X Genomics
• Immune repertoire / TCR / BCR sequencing
• Antibody discovery & engineering
• RNA-seq, Functional genomics, Differential expression
• Bioinformatics, Computational biology, Precision medicine
(full keyword list from earlier section applies here)
Pain Point Indicators
Companies experiencing:
• Coding bottlenecks in analysis workflows
• Bioinformatics dependencies slowing research
• Long analysis timelines (weeks ? hours opportunity)
• Integration challenges across tools
• Collaboration difficulties between wet lab and computational teams
Exclusion Criteria
Do not include:
• Pure software/IT companies (unless biotech-focused)
• Medical device companies (unless NGS-related)
• Clinical diagnostic labs (unless doing NGS research)
• Companies without active R&D programs
• Very early-stage startups (pre-Series A) without established research programs
???? Explicit Exclusion List — DO NOT include the following domains or organizations:
rosalind.bio
lifebit.ai
watershed.bio
latch.bio
kaist.ac.kr
yuhs.ac
helsinki.fi
Funding / Revenue Indicators
• Biotech: Series B+ funding or established revenue
• Pharma: Established companies with R&D budgets
• Academic / Research: Only major institutions actively publishing NGS research, but avoid those in the exclusion list above
Deliverables
• 1 CSV file for Next Generation Sequencing (companies only)
• Within the CSV: separate tabs/sheets (or separate CSVs if too large) for each application area or sub-industry
• Format: CSV with clearly labeled columns
Example CSV Structure
Companies
| Company Domain | Company Name | Industry | Sub-Industry / Application Area | LinkedIn | Employee Size | Location |
Key Notes
• Focus is NGS-related companies only (biotech, pharma, research)
• Employee size must be between 50–3,000 only
• Remote and distributed companies are also valuable
• ???? Companies/domains in the explicit exclusion list must be removed from the dataset
• Scale matters — bulk lists only (small datasets are not useful)
• Longer lists with mandatory fields > shorter lists with too many nice-to-haves