Python Developer (Web Scraping & Data Extraction)
QualityAI ·
- Category
- Software engineering
- Experience
- 3+ years
QualityAI ·
pwc · Bangalore (SDC) - Bagmane Tech Park
thinkbridge
Capital One · McLean, VA
Capital One · McLean, VA
Are you interested in working with the World’s leading AI-powered Quality Engineering Company? Ready to advance your career, team up with global thought leaders across industries and make a difference every day? Join us at QualityAI!
We are looking for a Python Developer (Web Scraping & Data Extraction) to join our growing team in United States
Location: US/CAN/ARG/MEX
About the Role
We're looking for an experienced Python developer who is strong in web scraping and data extraction to join our data engineering team. You'll build and maintain reliable, scalable pipelines that collect data from websites, APIs, and other online sources. This data feeds analytics, research, and decision-making across our business.
This role suits someone who has outgrown writing one-off scripts. You like building resilient, well-tested systems and take pride in data quality. Experience in finance, fintech, or investments is a real advantage, since much of our data supports market, investment, and financial analysis.
Key Responsibilities
Design, develop, and maintain scalable web scraping and crawling solutions in Python
Extract data from static pages, JavaScript-heavy sites, and REST/GraphQL APIs
Build robust pipelines that handle anti-bot measures, rate limits, CAPTCHAs, dynamic content, and site structure changes
Clean, validate, normalize, and transform raw data into analysis-ready datasets
Monitor scraper health, set up alerting, and quickly fix breakages to keep data fresh and accurate
Optimize scrapers for speed, reliability, and cost, including proxy management and concurrency
Store and manage data in relational and non-relational databases and cloud storage
Write clean, tested, well-documented code and take part in code reviews
Work with analysts, data scientists, and business stakeholders to turn data needs into technical solutions
Ensure scraping practices comply with terms of service, robots.txt, and data privacy regulations
Mentor junior developers and help set team standards for scraping and data quality
Required Qualifications
5-8 years of professional Python development experience, including at least 3 years focused on web scraping or data extraction
Strong hands-on experience with Scrapy, BeautifulSoup, Requests/HTTPX, and Selenium, Playwright, or Puppeteer
Experience with asynchronous and concurrent programming (asyncio, multithreading, multiprocessing)
Solid understanding of HTTP, HTML/CSS, DOM structure, XPath/CSS selectors, cookies, sessions, and authentication flows
Experience with proxies, user-agent rotation, headless browsers, and techniques for avoiding detection
Proficiency in data processing with Pandas, NumPy, and regular expressions
Working knowledge of SQL and at least one NoSQL database (e.g., PostgreSQL, MongoDB)
Experience building and scheduling data pipelines (e.g., Airflow, Prefect, or cron-based orchestration)
Familiarity with Git, CI/CD, Docker, and automated testing (pytest)
Experience with a cloud platform (AWS, Azure, or GCP)
Strong debugging and problem-solving skills, and clear written and verbal communication
Bachelor's degree in Computer Science, Engineering, Mathematics, Finance, or a related field (or equivalent experience)
Preferred Qualifications (Finance / Fintech / Investments)
Background in financial services, fintech, asset management, or investment research
Experience collecting and structuring data such as market prices, company filings (SEC/EDGAR), earnings reports, news, ESG data, alternative data, or fund and security reference data
Familiarity with financial concepts such as equities, fixed income, funds, portfolios, and risk and return metrics
Understanding of data quality, lineage, and audit requirements in regulated environments
Awareness of financial data licensing, compliance, and data governance considerations
Experience with financial data providers and APIs (e.g., Bloomberg, Refinitiv, Alpha Vantage, Yahoo Finance, Polygon)
Knowledge of NLP or text extraction from PDFs, filings, and unstructured documents
Experience with big data tools (Spark, Kafka) or data warehouses (Snowflake, BigQuery, Redshift)
Exposure to Kubernetes, infrastructure as code, or distributed scraping architectures
Key Competencies
Reliability mindset: builds scrapers that keep running and recover gracefully when sites change
Data quality focus: treats accuracy and completeness as non-negotiable
Ownership: takes responsibility for solutions end to end, from design to monitoring
Collaboration: communicates clearly with technical and non-technical stakeholders
Ethical judgment: respects legal, compliance, and data privacy boundaries
Benefits
Why QualityAI?
Capital One · McLean, VA