MarginaliaSearch

mirror of https://github.com/MarginaliaSearch/MarginaliaSearch.git synced 2025-02-23 21:18:58 +00:00

Author	SHA1	Message	Date
Viktor Lofgren	59a8ea60f7	(search) Further reduce the number of db queries by adding more caching to DbDomainQueries.	2025-01-10 14:15:22 +01:00
Viktor Lofgren	2d17233366	(search) Reduce the number of db queries a bit by caching data that doesn't change too often	2025-01-10 13:53:56 +01:00
Viktor Lofgren	6614d05bdf	(db) Make db pool size configurable	2025-01-09 20:20:51 +01:00
Viktor Lofgren	983d6d067c	(search-service) Add indexing indicator to sibling domains listing	2025-01-08 12:58:34 +01:00
Viktor Lofgren	d2864c13ec	(query-params) Add additional permitted query params	2025-01-07 20:21:44 +01:00
Viktor Lofgren	87d1c89701	(search) Add listing of sibling subdomains to site overview	2025-01-06 20:17:36 +01:00
Viktor Lofgren	8c69dc31b8	Merge branch 'master' into serp-redesign	2025-01-05 18:52:51 +01:00
Viktor Lofgren	4da3563d8a	(service) Clean up exceptions when requestScreengrab is not available	2025-01-04 14:45:51 +01:00
Viktor Lofgren	48d0a3089a	(service) Improve logging around grpc This change adds a marker for the gRPC-specific logging, as well as improves the clarity and meaningfulness of the log messages.	2025-01-02 20:40:53 +01:00
Viktor Lofgren	06efb5abfc	Merge branch 'master' into serp-redesign	2025-01-02 18:42:12 +01:00
Viktor Lofgren	78eb1417a7	(service) Only block on SingleNodeChannelPool creation in QueryClient The code was always blocking for up to 5s while waiting for the remote end to become available, meaning some services would stall for several seconds on start-up for no sensible reason. This should make most services start faster as a result.	2025-01-02 18:42:01 +01:00
Viktor Lofgren	8b05c788fd	(Search) Enable gzip compression of responses	2025-01-01 18:34:42 +01:00
Viktor Lofgren	ab5c30ad51	(search) Fix site info view for completely unknown domains Also correct the DbDomainQueries.getDomainId so that it throws NoSuchElementException when domain id is missing, and not UncheckedExecutionException via Cache.	2025-01-01 16:29:01 +01:00
Viktor Lofgren	2b222efa75	Merge branch 'master' into serp-redesign	2024-12-25 14:22:42 +01:00
Viktor Lofgren	1118657ffd	(system) Supply local IP to service discovery if multiFace is enabled	2024-12-19 22:20:19 +01:00
Viktor Lofgren	b1f970152d	(system) To support configurations with multiple docker networks, bind to the "most local" interface. Make the behavior optional.	2024-12-19 20:26:31 +01:00
Viktor Lofgren	e1783891ab	(system) To support configurations with multiple docker networks, bind to the "most local" interface.	2024-12-19 20:18:57 +01:00
Viktor Lofgren	0ce2ba9ad9	(jooby) Fix asset handler	2024-12-11 14:38:04 +01:00
Viktor Lofgren	b91463383e	(jooby) Clean up initialization process	2024-12-11 14:33:18 +01:00
Viktor Lofgren	fdee07048d	(search) Remove Spark and migrate to Jooby for the search service	2024-12-10 19:13:13 +01:00
Viktor Lofgren	fdc3efa250	(setup) Remove OpenNLP tokenization model This update eliminates all occurrences of the OpenNLP token model from the setup script, configuration, and test files, as this model file is no longer used.	2024-11-28 16:03:05 +01:00
Viktor Lofgren	52bc0272f8	(atag) Add alias domain support and improve domain handling Introduced optional alias domain functionality in EdgeDomain class to handle domain variations such as "www" in the anchor tags code, as there are commonly a number of relevant but glancing misses in the atags data.	2024-11-27 14:26:44 +01:00
Viktor Lofgren	14519294d2	Merge branch 'master' into live-search	2024-11-21 16:00:20 +01:00
Viktor Lofgren	51e46ad2b0	(refac) Move export tasks to a process and clean up process initialization for all ProcessMainClass descendents Since some of the export tasks have been memory hungry, sometimes killing the executor-services, they've been moved to a separate process that can be given a larger Xmx. While doing this, the ProcessMainClass was given utilities for the boilerplate surrounding receiving mq requests and responding to them, some effort was also put toward making the process boot process a bit more uniform. It's still a bit heterogeneous between different processes, but a bit less so for now.	2024-11-21 16:00:09 +01:00
Viktor Lofgren	47dfbacb00	(conf) Introduce a new concept of node profiles Node profiles decide which actors are started, and which views are available in the control GUI. This helps keep the system organized, and hides real-time clutter from the batch-oriented nodes.	2024-11-20 18:15:22 +01:00
Viktor Lofgren	a91ab4c203	(live-crawler) Crude first-try process for live crawling #WIP Some refactoring is still needed, but an dummy actor is in place and a process that crawls URLs from the livecapture service's RSS endpoints; that makes it all the way to being indexable.	2024-11-19 19:35:01 +01:00
Viktor Lofgren	6a3079a167	(search) Fix missing getter for proto	2024-11-18 21:05:22 +01:00
Viktor Lofgren	3791ea1e18	(service) Add a new application service for external liveness monitoring The new service 'status-service' will poll public endpoints periodically, and publish a basic read-only UI with the results, as well as publish the results to prometheus.	2024-11-17 18:01:08 +01:00
Viktor Lofgren	e5db3f11e1	(chore) Clean up some of the uglier delomboking artifacts	2024-11-15 13:57:20 +01:00
Viktor Lofgren	9f47ce8d15	(chore) Remove lombok There are likely some instances of delombok gore with this commit.	2024-11-11 21:14:38 +01:00
Viktor Lofgren	a5b4951f23	(chore) Remove use of deprecated STR.-style string templates	2024-11-11 18:02:28 +01:00
Viktor Lofgren	8b8bf0748f	(feature-extraction) Add new DocumentHeaders class encapsulating Html headers. Also adds a few new html features for CDNs and S3 hosting for use in ranking and query refinement.	2024-11-11 13:26:15 +01:00
Viktor Lofgren	bfeb9a4538	(feeds) Retire feedlot the feed bot, move RSS capture into the live-capture service	2024-11-09 17:56:43 +01:00
Viktor Lofgren	76e9053dd0	(setup) Move some file-downloads from setup script to the first boot of the control node of the system We can only do this for files that are not required for unit tests. As it is illegal to run more than one instance of the control service, this should be fine with regard to race conditions. The boot orchestration will also ensure that no other services will boot up before the downloading is complete.	2024-11-06 15:28:20 +01:00
Viktor Lofgren	45d3e6aa71	(download-sample) Break apart actor for better error recovery Change also adds logged events to give more feedback that something is happening.	2024-10-04 13:19:09 +02:00
Viktor Lofgren	d84a2c183f	(*) Remove the crawl spec abstraction The crawl spec abstraction was used to upload lists of domains into the system for future crawling. This was fairly clunky, and it was difficult to understand what was going to be crawled. Since a while back, a new domains listing view has been added to the control view that allows direct access to the domains table. This is much preferred and means the operator can directly manage domains without specs. This commit removes the crawl spec abstraction from the code, and changes the GUI to direct to the domains list instead.	2024-10-03 13:41:17 +02:00
Viktor Lofgren	23cce0c78a	Add a new function 'Live Capture' for on-demand screenshot capture The screenshots are requested by the site-service, and triggered via the site-info view.	2024-09-27 13:46:34 +02:00
Viktor Lofgren	1bd29a586c	(service-discovery) Add common base interface to all Grpc services To be able to tell service discovery whether to enable a service on a particular runtime, a common base interface DiscoverableService extends BindableService was added.	2024-09-27 13:46:34 +02:00
Viktor Lofgren	8047e77757	(doc) Correct dead links and stale information in the docs	2024-09-13 11:01:05 +02:00
Viktor Lofgren	abab5bdc8a	(index, EXPERIMENTAL) Evaluate using Varint instead of GCS for position data	2024-08-26 14:20:39 +02:00
Viktor Lofgren	b09e2dbeb7	(build) Fix dependency churn from testcontainers Apparently you need to pull in commons-codec now in order to run testcontainers, through spooky action at a distance.	2024-08-25 10:35:48 +02:00
Viktor Lofgren	0999f07320	(search-query) Add new ranking parameters for proximity and verbatim matches	2024-08-25 10:34:12 +02:00
Viktor Lofgren	0a383a712d	(qdebug) Accurately display positions when intersecting with spans	2024-08-15 11:44:17 +02:00
Viktor Lofgren	2ad93ad41a	(*) Clean up	2024-08-14 11:43:45 +02:00
Viktor Lofgren	34703da144	(slop) Support for nested array types and array-of-object types Also adding very basic support for filtered reads via SlopTable. This is probably not a final design.	2024-07-29 14:00:43 +02:00
Viktor Lofgren	aebb2652e8	(wip) Extract and encode spans data Refactoring keyword extraction to extract spans information. Modifying the intermediate storage of converted data to use the new slop library, which is allows for easier storage of ad-hoc binary data like spans and positions. This is a bit of a katamari damacy commit that ended up dragging along a bunch of other fairly tangentially related changes that are hard to break out into separate commits after the fact. Will push as-is to get back to being able to do more isolated work.	2024-07-27 11:44:13 +02:00
Viktor Lofgren	d36055a2d0	(keyword-extractor) Retire TfIdfHigh WordFlag This will bring the word flags count down to 8, and let us pack every value in a byte.	2024-07-17 13:54:39 +02:00
Viktor Lofgren	ae87e41cec	(index) Fix rare BitReader.takeWhileZero bug Fix rare bug where the takeWhileZero method would fail to repopulate the underlying buffer. This caused intermittent de-compression errors if takeWhileZero happened at a 64 bit boundary while the underlying buffer was empty. The change also alters how sequence-lengths are encoded, to more consistently use the getGamma method instead of adding special significance to a zero first byte. Finally, assertions are added checking the invariants of the gamma and delta coding logic as well as UrlIdCodec to earlier detect issues.	2024-07-16 11:03:56 +02:00
Viktor Lofgren	dfd19b5eb9	(index) Reduce the number of abstractions around result ranking The change also restructures the internal API a bit, moving resultsFromDomain from RpcRawResultItem into RpcDecoratedResultItem, as the previous order was driving complexity in the code that generates these objects, and the consumer side of things puts all this data in the same object regardless.	2024-07-16 08:18:54 +02:00
Viktor Lofgren	12590d3449	(index-reverse) Added compression to priority index The priority index documents file can be trivially compressed to a large degree. Compression schema: ``` 00b -> diff docord (E gamma) 01b -> diff domainid (E delta) + (1 + docord) (E delta) 10b -> rank (E gamma) + domainid,docord (raw) 11b -> 30 bit size header, followed by 1 raw doc id (61 bits) ```	2024-07-11 16:13:23 +02:00

1 2 3 4 5 ...

347 Commits