Senior Software Engineer, Data Acquisition

Remote Full-time
Note for all engineering roles: with the rise of fake applicants and AI-enabled candidate fraud, we have built in additional measures throughout the process to identify such candidates and remove them.About UsPeople Data Labs (PDL) is the provider of people and company data. We do the heavy lifting of data collection and standardization so our customers can focus on building and scaling innovative, compliant data solutions. Our sole focus is on building the best data available by integrating thousands of compliantly sourced datasets into a single, developer-friendly source of truth. Leading companies across the world use PDL’s workforce data to enrich recruiting platforms, power AI models, create custom audiences, and more.We are looking for individuals who can balance extreme ownership with a “one-team, one-dream” mindset. Our customers are trying to solve complex problems, and we only help them achieve their goals as a team. Our Data Engineering Acquisition Team ensures our customers have standardized and high quality data to build upon. You will be crucial in accelerating our efforts to build standalone data products that enable data teams and independent developers to create innovative solutions at massive scale. In this role, you will be working with a team to continuously improve our existing datasets as well as pursuing new ones. If you are looking to be part of a team discovering the next frontier of data-as-a-service (DaaS) with a high level of autonomy and opportunity for direct contributions, this might be the role for you. We like our engineers to be thoughtful, quirky, and willing to fearlessly try new things. Failure is embraced at PDL as long as we continue to learn and grow from it.What You Get to DoUse and develop web crawling technologies to capture and catalog data on the internetSupport and improve our web crawling infrastructureStructure, define, and model captured data, providing semantic data definition and automate data quality monitoring for data that we crawlDevelop new techniques to increase speed, efficiency, scalability, and reliability of web crawlsUse big data processing platform to build data pipelines, publish data, and ensure the reliable availability of data that we crawlWork with our data product and engineering team to design and implement new data products with captured data, and enhance and improve upon existing productsThe Technical Chops You’ll Need7+ years industry experience with clear examples of strategic technical problem solving and implementationStrong software development architecture and fundamentals for backend applicationsSolid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request/response)Solid programming experience: strong grasp of object-oriented design and experience building applications using asynchronous programming paradigms (e.g., async/await, event loops, or concurrency libraries)Experience building crawlersProficient in Linux / Unix command line utilities, Linux system administration, architecture, and resource managementExperience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)People Thrive Here Who CanMust thrive in a fast paced environment and be able to work independentlyCan work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)Strong written communication skills on Slack/Chat and in documentsYou are experienced in writing data design docs (pipeline design, dataflow, schema design)You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholdersSome Nice To HavesDegree in a quantitative discipline such as computer science, mathematics, statistics, or engineeringExperience as a Red TeamerExperience working in data acquisitionExperience in network architecture and how to debug and inspect network traffic (DNS, IPv4, Proxies, Application ports and interfaces; packet capture and analysis)Experience with Apache SparkExperience with SQL, including writing advanced queries (e.g., window functions, CTEs)Experience with streaming data platforms (e.g. Kafka or other pub/sub; Spark streaming or other stream processing)Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)Our BenefitsStockCompetitive SalariesUnlimited paid time offMedical, dental, vision insurance Health, fitness, and office stipendsThe permanent ability to work wherever and however you wantComp: $160K - $200KPeople Data Labs does not discriminate on the basis of race, sex, color, religion, age, national origin, marital status, disability, veteran status, genetic information, sexual orientation, gender identity or any other reason prohibited by law in provision of employment opportunities and benefits.Qualified Applicants with arrest or conviction records will be considered for Employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.Personal Privacy Policy for California Residentshttps://www.peopledatalabs.com/pdf/privacy-policy-and-notice.pdfOriginally posted on Himalayas

Apply Now
Apply Now →

Similar Jobs

Experienced Registered Behavior Technician for In-Home ABA Therapy - Atlanta, GA

Remote

Immediate Hiring: Experienced Registered Behavioral Technician (RBT) for Clinic-Based ABA Therapy Services

Remote

Experienced Registered Behavioral Technician (RBT) - ABA Therapy for Children with Autism Spectrum Disorder

Remote

Experienced Registered Nurse - Telehealth: Providing Remote Care Coordination and Patient Support

Remote

Experienced Substitute Teacher for Riverside County Schools - Join Scoot Education's Innovative Team

Remote

Experienced Substitute Teacher for San Bernardino County - Flexible Schedules & Competitive Pay

Remote

Experienced School Year Instructional Coach for High-Dosage Tutoring Programs in Edgewater Park, NJ

Remote

Experienced School Year Tutor for K-8 Students in Math and Literacy - Mickleton, NJ

Remote

Experienced Secondary Social Studies Teacher for Kansas - Flexible Hybrid Remote Arrangement

Remote

USPS Office Helper

Remote

**Experienced Customer Service Representative – Remote Customer Support**

Remote

Experienced Full Stack Data Entry Specialist – Remote Work Opportunity with blithequark

Remote

Urgently Require (USA) Coach/Ops Mgr Trainee in Charlotte, NC

Remote

Remote Admin Support - Data Entry Role

Remote

Looking for Online English Teacher (100% Remote) in Mesa, AZ

Remote

Urgently Need Corporate Fleet Coordinator- Remote Opportunity in Ann Arbor, MI

Remote

Motion Graphics Designer

Remote

Guest Service Representative - Overnight

Remote

Experienced Customer Experience Advocate – Enhancing Homebuyer Journeys at careerzynith

Remote

**Experienced Proofreader & Customer Representative Specialist – Remote – (DAY OR NIGHT SHIFT) in arenaflex**

Remote
← Back