Data Music Crawler — Multi-Source Data Acquisition

Freelance Software Engineer @ Freelance / Technical Consulting

Situation

The client needed reliable automated collection of music information from multiple external websites. Each source could have different:

  • Page structures.
  • Data formats.
  • Extraction rules.
  • Anti-bot behavior.
  • XPath requirements.
  • Data-processing requirements.

A single rigid crawler implementation would therefore be difficult to maintain. The system needed to dynamically select the appropriate crawler for each source and support flexible extraction rules.

Task

Build and maintain an automated music-data collection platform that gathers structured music information from multiple external sources, including Japanese radio/music stations and other music-data platforms. The system was designed to automate external data collection and prepare structured datasets for research, analytics, and business use.

Solution

Developed an automated multi-source crawling and data-processing system using PHP, Laravel, Selenium, Queue Jobs, MySQL, AWS, Linux, and Docker. The solution supported:

  • Multiple independent data sources.
  • Source-specific crawler assignment.
  • Dynamic and configurable extraction rules.
  • Flexible XPath-based extraction.
  • Selenium-based browser automation.
  • Queue-based asynchronous processing.
  • Rate-limit and anti-bot management.
  • ETL / ELT data-processing workflows.
  • External music-data sources such as Discogs and YouTube.
  • Production deployment and ongoing maintenance.

Contribution

As Freelance Software Engineer, I handled the technical lifecycle of the crawling platform. I:

  • Researched source-specific requirements and crawling technologies.
  • Designed the crawler architecture.
  • Developed source-specific crawler implementations.
  • Designed automatic crawler assignment based on data source.
  • Implemented Selenium-based crawling for sources requiring browser automation.
  • Developed Laravel Queue Jobs for asynchronous crawling and processing.
  • Implemented flexible XPath configuration.
  • Supported dynamic XPath extraction through configurable commands.
  • Integrated external music-data sources such as Discogs and YouTube.
  • Implemented data-processing pipelines.
  • Designed rate-limit and anti-bot handling approaches.
  • Deployed and maintained the production system.
  • Communicated technical requirements and decisions through the bridge PM.

Technologies

  • Languages: PHP
  • Frameworks / Libraries: Laravel, Selenium
  • Data Acquisition: Web Crawling, Web Scraping, Dynamic XPath Extraction, Browser Automation
  • Data Processing: Laravel Queue Jobs, ETL, ELT, Data Transformation, Data Normalization
  • External Data Sources: Discogs, YouTube, Japanese music/radio sources
  • Database / Storage: MySQL
  • Cloud / Infrastructure: Linux, AWS
  • DevOps: Git, Docker

Results

  • Automated music-data collection across multiple external sources.
  • Supported source-specific crawler assignment.
  • Implemented flexible XPath-based data extraction.
  • Automated asynchronous crawling and processing through queue jobs.
  • Supported Selenium-based browser automation.
  • Integrated external music-data sources including Discogs and YouTube.
  • Established ETL/ELT processing workflows.
  • Maintained the crawler platform in production.

Project Details

Period

2020 – 2021

Domain

Data Acquisition / Web Crawling / Automation / ETL

Tech Stack
PHPLaravelSeleniumMySQLLaravel Queue JobsETL / ELTAWSDocker
Tags
#data-engineering#web-crawling#etl#automation