Описание:
Vantage Point Global is a talent consultancy established in 2014 that supports top-tier clients in building their talent pipelines. It partners with clients to create career opportunities and deliver a future-ready workforce, providing continuous formal training, personal development, and mentoring. The company works with clients globally and has teams in London, Manchester, Newcastle, Belfast, Warsaw, and Wroclaw, with plans to expand into the US and Asia.
Задачи:
- Architect, develop, and maintain high-throughput data pipelines in Databricks and AWS;
- Ingest, normalize, and enrich large volumes of structured and unstructured data from internal, market, vendor, and alternative sources;
- Collaborate with AI Engineers, ML scientists, and software teams to translate requirements into scalable data architectures, schemas, and APIs;
- Optimize pipeline performance and cost using distributed processing techniques and AWS best practices;
- Enforce data governance, privacy, and lineage standards;
- Catalogue assets in Unity Catalog and manage PII/PCI classification;
- Build automated validation, testing, and monitoring frameworks for data quality and freshness across offline and online workloads;
- Support onboarding and integration of external data vendors, ensuring compliance and rapid time-to-value;
- Evaluate emerging GenAI tooling and drive proof-of-concepts;
- Own tools and workflows for compliant web crawling and scraping;
- Test and validate scraped data for accuracy, quality, and compliance;
- Identify and resolve scraping issues and scale processes as needed.
Требования:
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field;
- 5+ Years of experience in data engineering or a related role;
- Strong Python and SQL skills;
- Experience with Spark or Scala and distributed data processing;
- Proficiency in building scalable, distributed data pipelines in a cloud environment;
- Familiarity with Linux/UNIX, HTTP, HTML, JavaScript, and networking concepts;
- Knowledge of web scraping tools and libraries, including Requests, BeautifulSoup, Scrapy, Pandas, Selenium, or Spark;
- Working knowledge of version control systems and open-source practices;
- Solid understanding of data architecture, data modeling, and data warehousing;
- Excellent analytical and problem-solving skills;
- Strong written and spoken English skills;
- Commitment to the highest ethical standards;
- Nice to have: extracting text from PDFs, images, and applications;
- System monitoring or administration tools;
- Graph databases;
- Analysing big data sets.
Условия:
The application requires a CV and answers to a few initial questions; Applicants who do not hear back within three weeks of applying will not progress.