Webscraping Manufacturing and Biopharma Company Domains

Please login or register as jobseeker to apply for this job.

TYPE OF WORK

Full Time

WAGE / SALARY

Php 1500

HOURS PER WEEK

TBD

DATE UPDATED

Oct 2, 2025

JOB OVERVIEW

Project Brief (Companies Only – Next Generation Sequencing Focus)

Important Note (Read First)
For this project, longer lists with mandatory attributes are more valuable than shorter lists with many nice-to-haves.
If you must choose, always prioritize volume with mandatory fields over completeness with nice-to-haves.

Accuracy still matters — no duplicates, no fake LinkedIns, no broken domains.

Nice-to-have attributes are welcome, but they should not reduce the overall scale of the dataset.


What We Need Collected

Companies

Mandatory:
• Company domain (website URL) OR
• Company name

Nice-to-have (if available):
• Company name
• Industry / sub-industry classification
• Location (city, state, country)
• Employee size (only include companies with 50–3,000 employees, reported in ranges:
• 51–200
• 201–500
• 501–1,000
• 1,001–3,000)
---------- pany page


Volume & Expectations
• NGS-focused dataset: minimum goal of 5,000 companies globally, with emphasis on:
• Primary regions: North America, Europe
• Secondary regions: Australia, Israel, Singapore
• Within the dataset, create sub-lists by research/application area (e.g., oncology research, immunotherapy, vaccine development, cell therapy, biomarker discovery).

???? If you cannot fully reach the target, you must report blockers (e.g., lack of data sources, technical limits) in a timely manner.


Data Quality
• Companies must be within 50–3,000 employees only
• No duplicates, no broken domains, no fake LinkedIn profiles
• Prioritize Company Domain whenever available
• Must clearly fit into NGS ecosystem (see keyword framework and exclusion rules below)


Target Company Indicators

Key Search Keywords
• Primary NGS: Next-generation sequencing, Genomics, Transcriptomics, Single-cell RNA-seq, 10X Genomics
• Immune repertoire / TCR / BCR sequencing
• Antibody discovery & engineering
• RNA-seq, Functional genomics, Differential expression
• Bioinformatics, Computational biology, Precision medicine

(full keyword list from earlier section applies here)


Pain Point Indicators

Companies experiencing:
• Coding bottlenecks in analysis workflows
• Bioinformatics dependencies slowing research
• Long analysis timelines (weeks ? hours opportunity)
• Integration challenges across tools
• Collaboration difficulties between wet lab and computational teams



Exclusion Criteria

Do not include:
• Pure software/IT companies (unless biotech-focused)
• Medical device companies (unless NGS-related)
• Clinical diagnostic labs (unless doing NGS research)
• Companies without active R&D programs
• Very early-stage startups (pre-Series A) without established research programs

???? Explicit Exclusion List — DO NOT include the following domains or organizations:

----------
----------
----------
----------
----------
----------
----------
----------
----------
----------
rosalind.bio
----------
lifebit.ai
----------
watershed.bio
latch.bio

----------
----------
---------- t
----------
----------
----------
----------
----------
kaist.ac.kr
----------
----------
yuhs.ac
----------
----------
helsinki.fi
----------
----------



Funding / Revenue Indicators
• Biotech: Series B+ funding or established revenue
• Pharma: Established companies with R&D budgets
• Academic / Research: Only major institutions actively publishing NGS research, but avoid those in the exclusion list above


Deliverables
• 1 CSV file for Next Generation Sequencing (companies only)
• Within the CSV: separate tabs/sheets (or separate CSVs if too large) for each application area or sub-industry
• Format: CSV with clearly labeled columns


Example CSV Structure

Companies
| Company Domain | Company Name | Industry | Sub-Industry / Application Area | LinkedIn | Employee Size | Location |


Key Notes
• Focus is NGS-related companies only (biotech, pharma, research)
• Employee size must be between 50–3,000 only
• Remote and distributed companies are also valuable
• ???? Companies/domains in the explicit exclusion list must be removed from the dataset
• Scale matters — bulk lists only (small datasets are not useful)
• Longer lists with mandatory fields > shorter lists with too many nice-to-haves

VIEW OTHER JOB POSTS FROM:
SHARE THIS POST
facebook linkedin