MarginaliaSearch/code/features-crawl/readme.md

9 lines
403 B
Markdown
Raw Normal View History

# Crawl Features
These are bits of search-engine related code that are relatively isolated pieces of business logic,
that benefit from the clarity of being kept separate from the rest of the crawling code.
2024-02-06 15:29:55 +00:00
* [content-type](content-type/) - Content Type identification
* [crawl-blocklist](crawl-blocklist/) - IP and URL blocklists
* [link-parser](link-parser/) - Code for parsing and normalizing links