Data Twitter Crawler — Distributed Data Collection
Freelance Software Engineer @ Freelance / Technical Consulting
Situation
The client needed automated collection of Twitter/X data to analyze production feedback and user activity. The system had to handle:
- Large volumes of tweets and related social data.
- Long-running crawling jobs.
- Multiple crawlers executing in parallel.
- Distributed execution across multiple VPS instances.
- Twitter/X API rate limits.
- Shared account/API constraints.
- Avoiding overlapping crawling work.
- Storage and export of millions of records.
- Reliable processing and monitoring of long-running jobs.
Task
Build and maintain an automated Twitter/X data collection system capable of continuously collecting large volumes of social-media data for production feedback analysis and data analytics. The system needed to collect and process data at scale while distributing crawler workloads across multiple VPS environments and managing rate limits for shared Twitter/X accounts.
Solution
Developed a distributed multi-crawler data acquisition and processing system using PHP, Laravel, Queue Jobs, Horizon, MySQL, AWS, Linux, Docker, and Ansible. The system used a master/child job model to split large crawling tasks into smaller units that could execute in parallel across multiple VPS environments. Core capabilities included:
- Master job orchestration.
- Child job generation and workload distribution.
- Parallel crawler execution.
- Multi-VPS deployment.
- Queue-based asynchronous processing.
- Laravel Horizon job management.
- Workload balancing across crawler instances.
- Twitter/X rate-limit management.
- Large-scale MySQL data storage.
- Data export for analytics.
- Long-running job execution.
- Automated deployment and infrastructure configuration using Ansible.
Contribution
As Freelance Software Engineer, I was responsible for the technical lifecycle of the crawling system. I:
- Researched crawling and Twitter/X data-collection requirements.
- Designed the distributed crawler architecture.
- Designed master/child job orchestration.
- Implemented parallel crawling across multiple VPS environments.
- Implemented Laravel Queue Jobs.
- Used Laravel Horizon for queue monitoring and management.
- Designed workload balancing between crawler instances.
- Implemented rate-limit management.
- Designed storage for millions of data records.
- Implemented data processing and export functionality.
- Configured Linux/AWS infrastructure.
- Automated infrastructure/deployment tasks using Ansible.
- Maintained Docker-based environments.
- Monitored and maintained long-running production jobs.
- Communicated technical decisions and requirements through the bridge PM.
Technologies
- Languages: PHP
- Frameworks / Libraries: Laravel, Laravel Queue, Laravel Horizon, Blade, Bootstrap
- Data Acquisition: Twitter/X data crawling, Web crawling, Data scraping
- Data Processing: ETL, ELT, Queue-based processing, Large-scale data export
- Database / Storage: MySQL
- Cloud / Infrastructure: Linux, AWS, Multiple VPS environments
- DevOps / Automation: Git, Docker, Ansible, Laravel Horizon
Results
- Automated Twitter/X data collection at production scale.
- Distributed crawling workloads across multiple VPS environments.
- Processed and stored millions of rows of Twitter/X-related data.
- Collected data including tweets, mentions, followers, and followings.
- Implemented parallel crawler execution.
- Established master/child job orchestration.
- Implemented long-running asynchronous processing.
- Added rate-limit management for shared API/account constraints.
- Provided data export capabilities for analytics use.
Project Details
2020 – 2021
Data Acquisition / Web Crawling / Automation / Analytics