The firehose is real
Here is the uncomfortable arithmetic. arXiv's machine-learning categories โ cs.LG,
cs.CL, cs.CV, stat.ML โ receive hundreds of new submissions every single day.The exact count fluctuates, but the cs.LG + cs.CL + cs.CV listings routinely run into the hundreds of new papers per weekday, and the trend has only gone up. No human reads a meaningful fraction of that. Even if you could read one paper carefully in thirty minutes, and you did nothing else for eight hours, you'd cover sixteen. The daily inflow is an order of magnitude past that โ before you sleep, eat, or do your actual job.
So the instinct kicks in: skim everything. Open the daily listing, read titles and abstracts, star the interesting ones, move on. This feels like keeping up. It is the single most seductive failure mode in the field.
The goal is not coverage. Coverage is impossible and chasing it is how you burn out. The goal is a working model of the field that's deep where it matters to you and shallow but current everywhere else โ and you get there by splitting those two jobs apart.
Read deeply vs monitor
These are different activities with different tools. Conflating them is most of the problem.
Read deeply (a few / week)
Papers central to your work or to a shift you need to understand. Read the whole thing, re-derive the key equation, check the ablations, note what they didn't test. This is where real understanding compounds. Budget 3โ5 of these a week, no more.
Monitor (everything else)
The broad current โ what's trending, what's contested, what just dropped. You want awareness, not mastery: enough to know a thing exists and roughly why it matters, so you can pull it into the "read deeply" bucket if it becomes relevant. Outsource this to curators.
A starter set of sources
Breadth is a solved problem if you let the right people solve it for you. Here's a small, honest list โ each with why it earns a slot and who it's not for. Pick three or four, not all eight; a bloated feed is just the firehose wearing a costume.
- arXiv
cs.CL/cs.LGdaily listings โ the raw source. Worth a scan (titles only) if your work is research-adjacent, to catch things before curators do. Not for you if you want signal over completeness โ the noise floor is brutal. - Semantic Scholar feeds & alerts โ follow specific authors, papers, or topics and get notified of new and citing work. Best for tracking a narrow area you actually care about. Not a discovery tool for things outside your existing interests.
- The Batch (DeepLearning.AI) โ Andrew Ng's weekly newsletter. Calm, well-edited, business-and-research balance. Great for monitoring without the hype. Too high-level if you want method-level depth.
- Latent Space (newsletter + podcast) โ sharp coverage of applied AI and the engineering reality behind the research. Best for practitioners shipping with these models. Less focused on pure theory.
- Yannic Kilcher (YouTube) โ long, opinionated paper walkthroughs that actually engage with the method and the weaknesses. Excellent for understanding a specific paper deeply. Not a breadth tool โ it's one paper at a time.
- AK / @_akhaliq (on X / Hugging Face) โ a firehose-but-curated stream of the day's notable papers, usually with the key figure attached. Great for monitoring what's breaking right now. Easy to doomscroll โ set a time box.
- Hugging Face Daily Papers โ a community-voted daily shortlist of papers, with discussion. A solid monitor layer that surfaces what others found worth reading. Skews toward what's popular, not necessarily what's important.
- Import AI (Jack Clark) โ weekly newsletter with a policy and big-picture lens. Valuable for situating research in the wider world. Not where you go for implementation detail.
How I triage a paper in five minutes
When something graduates from "monitor" to "maybe read deeply," you still need a fast gate before committing thirty minutes. Here's the routine.
The discipline is in step five. "Save for later" without a reason is how you build a graveyard of good intentions. One line on why this matters to me is the difference between a reading list and a museum of optimism.
Where this is going
The honest end state of all this is that curation is the bottleneck, not access. Everyone has the same arXiv. The edge is having a trusted filter tuned to your taste โ your sources, your topics, your bar for "worth it" โ that does the wide monitoring so you can spend your scarce attention reading deeply.
That's the whole idea behind Clove: you pick the sources and describe what you care about in plain English, and the fox reads the firehose so you don't have to โ sending only what clears your bar, to a place you already check. Keeping up with AI research is just one of the firehoses it was built to tame. If "monitor broadly, read deeply, skim nothing" sounds like the system you want but never have time to maintain, that's the gap it's meant to close.
Sources & further reading
- arXiv โ Computation and Language (cs.CL)arxiv.org
- Semantic Scholar โ research feeds & alertssemanticscholar.org
- The Batch โ DeepLearning.AIAndrew Ng ยท deeplearning.ai
- Latent Spacelatent.space
- Yannic Kilcher โ paper walkthroughsYouTube
- Hugging Face โ Daily Papershuggingface.co
- Import AIJack Clark ยท importai.net