-
Notifications
You must be signed in to change notification settings - Fork 0
feat: Parallel/concurrent collection #95
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Conversation
This commit introduces parallel collection capabilities to the `borg` CLI, significantly improving the performance of large-scale data collection. Key features and changes include: - **Parallel Downloads:** A `--parallel` flag has been added to the `collect github repos` and `collect website` commands, allowing users to specify the number of concurrent workers for downloading and processing. - **Rate Limiting:** A `--rate-limit` flag has been added to the `collect website` command to control the maximum number of requests per second to a single domain, preventing the crawler from overwhelming servers. - **Graceful Shutdown:** The worker pools now respect context cancellation, allowing for a graceful shutdown on interrupt (e.g., Ctrl+C). This improves the user experience for long-running collection tasks. - **Refactored Downloaders:** The `github` and `website` downloaders have been refactored to use a robust worker pool pattern, with proper synchronization primitives to ensure thread safety. Co-authored-by: Snider <631881+Snider@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
Warning Rate limit exceeded
⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ✨ Finishing touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Adds a placeholder file at `pkg/player/frontend/demo-track.smsg` to resolve the CI build failure caused by a missing asset required by Go's `embed` directive. Co-authored-by: Snider <631881+Snider@users.noreply.github.com>
Adds a placeholder file at `pkg/player/frontend/demo-track.smsg` to resolve the CI build failure caused by a missing asset required by Go's `embed` directive. Co-authored-by: Snider <631881+Snider@users.noreply.github.com>
This submission adds parallel collection capabilities to the
borgCLI, focusing on thecollect github reposandcollect websitecommands. It introduces a worker pool to manage concurrency, adds per-domain rate limiting for responsible crawling, and enhances the progress bar to reflect the status of parallel tasks. The implementation also includes graceful shutdown on interrupt.Fixes #28
PR created automatically by Jules for task 8563216625221556707 started by @Snider