PriEco: New open source web search engine

Hello! I have worked on the decentralization. Decentralization has the potential to solve many of the current PriEco problems.

You can now self-host PriEco, either compile binary from source code, run “docker compose up” on the code or using docker/podman. The easiest is docker compose up.

I’d really appreciate it if you could try it out and tell me if it works for you, overall simplicity of the set-up and if you’ll keep on self-hosting PriEco node.

Positives

Decentralization allows self-hosting of an independent web search engine (PriEco) as it gives nodes a clear way to get results for their users’ queries.

Improves trust and audit of the code as I wouldn’t be the only one running PriEco.

The nodes themselves are beneficial for the entire network as they fetch and parse web pages. Which means they improve crawling capacity of the entire PriEco network which could allow us to keep up with established web search engines.

Nodes work offline/disconnected as they save the results they fetched from the network

And makes government censorship that much harder as they may block https://prieco.net/ but blocking other nodes or even full access to the index would be bordering needing to block internet access altogether.

Negative

Nodes can serve results only when they know query and preferred language+location. This means sharing this data with potentially untrustworthy actors. The problem is that the query itself may contain sensitive information. If this behavior is undesirable and user trusts only the node they are connected to, such as https://prieco.net/ they can disable decentralization in settings.

How to self-host?

Binary

By the way, we are now mirrored to GitHub and GitLab. If this makes it easier for you to work with.

  1. Clone PriEco repository
  2. “cd” into the dir and run cargo run -r. You need to have rust and cargo installed

Compose

In the cloned repository edit docker-compose.yamlfile as it contains the settings. More in settings.

Then run docker compose up

Image

You can use prebuilt image

Full command with all settings would be:

`docker run -d
–name prieco
-p 127.0.0.1:8000:8000
–read-only
-t
–memory=8g
–memory-reservation=2g
–cpu-shares=1024
–blkio-weight=500
–tmpfs /tmp:size=1G
-v tantivy:/app/data/tantivy
-v blobs:/app/data/blobs
-v meta:/app/data/meta
-v vectors:/app/data/vectors
-v $(pwd)/models:/app/models:ro,z
-v $(pwd)/data/bge/model.onnx:/app/data/model.onnx:ro,z
-e PRIECO_IP=0.0.0.0
-e PRIECO_PORT=8000
-e PRIECO_TANTIVY=data/tantivy
-e PRIECO_META=data/meta
-e PRIECO_VECTOR=data/vectors
-e PRIECO_BLOB=data/blobs
-e PRIECO_TICKET=“”
-e PRIECO_WORKERS=0
-e PRIECO_MAX_STORAGE=0
-l io.containers.autoupdate=registry
-l com.centurylinklabs.watchtower.enable=true
–ulimit nofile=65535:65535
–restart unless-stopped
jojoyou/prieco:latest`

Maybe you’d want to run a bandwidth limiter too

`docker run -d
–name bandwidth-limiter
–network container:prieco
–cap-add NET_ADMIN
-e NETWORK_RATE=unlimited
–restart unless-stopped
alpine:latest
sh -c “apk add --no-cache iproute2 && if [ “$NETWORK_RATE” = “unlimited” ]; then tc qdisc del dev eth0 root 2>/dev/null || true; else tc qdisc add dev eth0 root tbf rate $NETWORK_RATE burst 128kbit latency 400ms; fi && sleep infinity”`

And definitely you’d want to run auto updater

`docker run -d
–name watchtower
–restart unless-stopped
-v /var/run/docker.sock:/var/run/docker.sock
nickfedor/watchtower
–label-enable
–cleanup
–interval 3600`

Settings

Most important are:
-p 127.0.0.1:8000:8000

  • Modify 127.0.0.1:8000 however you wish and PriEco is going to listen to this IP and Port

-e PRIECO_TICKET=“”

  • You need a ticket of a node which is already in the network
  • CONTACT ME FOR IT, later I’ll make it public and anyone from the network can invite you using their ticket

-e PRIECO_WORKERS=0

  • How many web pages you allow PriEco to fetch at the same time

Design

I use https://www.iroh.computer/ for decentralization

Once node starts, it prints its Ticket which other nodes can use to connect to and thus the whole network.

Your node creates its profile describing in what type of searches it is good at. And broadcasts it through the network.

When a user queries a node, it decides max 4 other nodes that might be good for this query too and calls them using iroh

User queried node, stores returned results in its index for local future retrieval

Index updater

I sometimes update the index itself. For example right now I’m running an update that pings domains and results in the index and removes dead ones, as it started to be a real problem.

To achieve this I’ve designed and created an updater mechanism. It easily allows me to create an update and share it between nodes. The nodes slowly apply it for their index in a way that doesn’t overload them nor slow down query time. And I can create multiple updates and a node will handle all of them, at once, in a very efficient way. It doesn’t even matter if 1 update is in progress and another one comes in.

Mission

I’m excited to have a decentralized web search engine with own index. There are and were a few options before but all of them are not great. I hope this dramatically progresses our mission with PriEco, and I’d really hope it allows us to crawl the web with huge speed making PriEco competitive with established search engines

Feedback

Decentralization still isn’t very tested. I’d really appreciate it if you could help me test it all out, improve it and fetch the web.

5 Likes