Hello! I’m new to Yacy and I am hosting an instance. The results have a lot of duplicates from the same domain with multiple pages in a category. How can I fix the results to display only 1 result per domain? or at least reduce duplicates
I set these two:
fuzzy_signature_unique_b:true^100000.0f
host_extent_i=true^100000.0f
Then hit:
Set Boost Query
It’s better, but still the same domain over and over
Hi, I built an experimental fork of YaCy to try fixes for weaknesses I measured on freeworld (in 16 queries, only 11% of the top-10 results contained all query terms). It is not meant to replace anything, just to share what the changes do, with data:
- Ranking: minimum match 2<-1 5<80% instead of mm=1, term coverage weighting after per-peer normalization, a thin-page weight, and searching all peers of small networks.
- CJK: bigram indexing in Solr and in the word index.
- Trust layer: Ed25519 peer keys, signed seeds, coordinator-signed trust lists with declared tags (e.g. ads), and per-document author signatures.
- NAT traversal: peers behind a NAT answer searches through a libp2p circuit relay (small go-libp2p sidecar).
In closed docker-compose networks (3 upstream vs 3 fork peers), R-precision went from 0.52 to 0.84–0.88 and keyword-stuffed pages in the top 5 from 14 to 0. The trust/NAT setup passes 26/26 checks. It is incompatible with freeworld by design, and the corpus is synthetic, so real-network numbers are still open.
Code: GitHub - pad01g/yacy_search_server: YaCy fork with search quality fixes: strict Solr minimum match, term coverage weighting, CJK bigram indexing. Experiment: pad01g/yacy-lab · GitHub (branch improved-search)
Lab and demo: GitHub - pad01g/yacy-lab: Closed-network docker compose experiment comparing upstream YaCy and pad01g/yacy_search_server (search quality fixes) · GitHub
I’d be glad to hear whether any of this is worth proposing upstream as small PRs, starting with the mm setting and the CJK word count. (Related: /t/3400.)
Disclaimer: everything is vibe-coded with Claude Code.