MarginaliaSearch/code/processes/crawling-process/java/nu/marginalia/crawl/retreival
2024-04-22 15:37:35 +02:00
..
fetcher (crawler) Strip W/-prefix from the etag when supplied as If-None-Match 2024-04-22 14:31:05 +02:00
revisit (crawler) Code quality 2024-04-22 15:37:35 +02:00
sitemap (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
Cookies.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
CrawlDataReference.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
CrawlDelayTimer.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
CrawledDocumentFactory.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
CrawlerRetreiver.java (crawler) Use the probe-result to reduce the likelihood of crawling both http and https 2024-04-22 15:36:43 +02:00
CrawlerWarcResynchronizer.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
DomainCrawlFrontier.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
DomainProber.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
LinkFilterSelector.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00
RateLimitException.java (refac) Remove src/main from all source code paths. 2024-02-23 16:13:40 +01:00