We build and operate large-scale web crawling and extraction systems that turn unstructured, dynamic websites into clean, structured, and correct data. You'll own crawlers
What they ask for
3+ years of professional software development (5+ for the senior track).
Strong Python - real production code, async.
Hands-on web scraping at scale, including firsthand experience with anti-bot
A data-correctness mindset - you care that extracted data is right, not just that the job finished.
Docker + at least one major cloud (GCP, Azure, or AWS), Git, comfortable on the CLI.
Clear communicator who works well in a small, fast-moving team.
More details and full description
Skills mentioned
Python
JavaScript
Work eligibility
Add your citizenship to check
Germany · EU · Guidance only, not legal advice. Check official sources for your situation.
Full description
About the Job
We build and operate large-scale web crawling and extraction systems that turn unstructured, dynamic websites into clean, structured, and correct data. You'll own crawlers
end-to-end: discovery, resilient fetching against real anti-bot defenses, and turning HTML/PDF into trustworthy structured output.
What you will do
• Build and scale crawlers that handle dynamic, large-scale sites - and keep them running as those sites change and fight back.
• Stay ahead of anti-bot measures: TLS/browser fingerprinting, proxies, session and rate-limit strategy, change detection.
• Turn raw HTML and PDF into clean, correct structured data - including LLMassisted extraction - where silently-wrong output is worse than a crash.
• Design and operate pipelines for ingestion, deduplication, versioning, and enrichment.
• Ship and run containerized apps in the cloud, with CI/CD.
• Shape our technical architecture and engineering culture.
Our Tech Stack
Python, Docker, GCP (Cloud Run jobs) and Azure Blob, MongoDB. Go services and a React frontend on the periphery. You don't need all of it on day one, but you do need to
ramp up fast
Your profile
• 3+ years of professional software development (5+ for the senior track).
• Strong Python - real production code, async.
• Hands-on web scraping at scale, including firsthand experience with anti-bot
defenses and the failure modes of large-scale collection.
• A data-correctness mindset - you care that extracted data is right, not just that the job finished.
• Docker + at least one major cloud (GCP, Azure, or AWS), Git, comfortable on the CLI.
• Clear communicator who works well in a small, fast-moving team.
Nice to have
• LLM-assisted extraction / working with LLMs, ML pipelines, or data labeling workflows.
• Go or Java/Spring Boot.
• Frontend (React) — or eagerness to pick it up.
• Terraform, Kubernetes, modern DevOps tooling.
Why us?
• A vibrant in-office culture in Munich with the option for up to 60 % work from home.
• Flexible working hours.
• 30 vacation days per year (including 4 “company rest days” over christmas).
• 30 days of "workation" per year, within the EU and selected countries.
• High autonomy and flat hierarchies.
• EGYM Wellpass for unlimited access to fitness courses and gyms.
• Udemy access for educational videos.
Closing
Certivity is an equal opportunities employer. We are committed to equal employment opportunities regardless of race, religion, sexual orientation, age, marital status, disability, or gender identity. Please do not submit personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, data concerning your health, or data concerning your sexual orientation.
Find more English Speaking Jobs in Germany on Arbeitnow
Für ein etabliertes Unternehmen suchen wir eine erfahrene SAP-Beratung für FI/CO, fest angestellt und im Haus. Du unterstützt die Fachbereiche bei der Weiterentwicklung der SAP-Prozesse und begleitest den Umstieg auf…
<p>At the heart of our offering is a versatile software platform for orchestrating complex robotic systems, from industrial robot arms and mobile manipulators to humanoids. Our Forward Deployed engineers bring it into…