TripleChoice Inc

Senior Web Scraping Engineer

Check with seller / month
Mumbai, Maharashtra, India IT Engineer & Developer Active
Actively Hiring Mumbai Full Time
Advertisement

Job Description


S
Seven Consultancy


Sales Engineer-Reputed Laptop Manufacturing Industry-Mumbai, Maharashtra, India-3 lakhs-Vishal
Seven Consultancy • Mumbai, Maharashtra • via BeBee
14 hours ago
₹9L–₹10.8L a year
Full–time
No Degree Mentioned
Apply on BeBee
Job description
Job Details
• Conduct customer need analysis and present possible solutions.
• Understand product functionality and potential problems, finding ways to handle them.
• Determine best products to meet customer requirements.
• Assist sales department with technical expertise on products or services.
• Participate in exhibitions, conventions, and other events to boost sales.
• Train colleagues and sal...
Show full description
Report this listing

TripleChoice Inc


Senior Web Scraping Engineer
TripleChoice Inc • Thane, Maharashtra • via BeBee
18 hours ago
₹40L–₹50L a year
Full–time
No Degree Mentioned
Apply on BeBee
Job description
Job Description

We are seeking a Senior Web Scraping Engineer to join our team. As a key member of our data ingestion pipeline, you will be responsible for designing and implementing an HTTP-first crawler with a Playwright fallback.
• Design an HTTP-first crawler (Scrapy or aiohttp) with Playwright fallback only for JS-heavy pages.
• Implement sitemap diffing and conditional GETs (ETag/Last-Modified) for incremental runs.
• Build a lightweight 'needs JS?' classifier (HTML length, JSON-LD presence, data-product markers) to auto-route HTTP vs Playwright.
• Enforce per-domain throttles/backoff (2–4 concurrent/domain; auto-lower on 429/503).
• Add URL normalization/canonicalization and de-dup (respect ; hash PDFs).
• Handle PDF discovery & download (HEAD first to dedupe; size/concurrency caps; SHA-256 keys).
• Apply Playwright browser automation resource budgets (block images/fonts/analytics; kill outliers by size/CPU/time).
• Integrate third-party APIs (REST/GraphQL) as first-class sources: handle auth (API keys/OAuth2), pagination, and rate limits; unify API + crawl outputs.
• Own automation & orchestration for scheduled runs (Airflow/Temporal/Celery or cron), idempotent retries, and alerting.
• Create per-domain selectors (YAML) with verification on hold-outs; re-learn only when health drops.
• Ship observability: per-site field coverage, error rates, retries, avg page time, and PDF success.
• Maintain allow/deny paths; adhere to robots.txt and Terms of Service.
• Containerize workers; provide runbooks/CI; collaborate with data team on schemas/normalization.

Must-have Qualifications:
• 4+ years Python, including 2+ years building production web crawlers at scale.
• Strong with Scrapy or aiohttp/asyncio and Playwright (or Puppeteer) in production.
• Practical proxy management, polite anti-bot tactics, and per-domain rate limiting.
• Hands-on with ETag/Last-Modified, retries, backoff, and HTTP caching.
• Confident with CSS/XPath, schema.org/JSON-LD, and HTML parsing.
Ready to take the next step?

Don't wait — new applications are being reviewed daily.

Job Safety Alert Real jobs on Jobsiya are always free. Never pay for an interview and never share bank or OTP details. Report this job →