- Python 70.7%
- HTML 19.2%
- CSS 9.1%
- Dockerfile 1%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| app | ||
| config | ||
| static/css | ||
| templates | ||
| .dockerignore | ||
| .gitignore | ||
| docker-compose.yml | ||
| Dockerfile | ||
| entrypoint.py | ||
| README.md | ||
| requirements.txt | ||
| todo.md | ||
Newsgetter
A personal news aggregator. It collects articles from RSS feeds (and similar sources), scores them with an LLM according to your interests, and surfaces the ones worth your attention.
Features
- RSS/Atom feed collection with deduplication
- LLM-based relevance scoring (0–100) with a short justification shown next to every article
- Topic tabs with categories assigned by the LLM
- A "Rejected" view of the last 7 days, so false negatives can be restored
- Periodic source checking (default: every 4 hours) with scoring right after every fetch
- Automatic retention cleanup (rejected after 7 days, read after 21 — both configurable)
- Application log with rotation and an in-GUI viewer
- Single SQLite database, no external services required
Quick start
The image is published on Docker Hub as qbsoon/newsgetter:latest:
docker compose up -d
The web interface is available at http://localhost:8000.
To build the image from source instead:
docker build -t qbsoon/newsgetter:latest .
The container starts as root and immediately drops privileges to whoever can write the mounted data directory, so any bind-mount just works — no host-side chown or compose changes needed:
- Mount owned by a regular user → the app runs as that user.
- Root-owned mount (e.g. TrueNAS datasets,
root:rootwith modify rights for theappsuser via ACLs) → the app probes well-known unprivileged identities (apps 568:568, user 1000:1000, the mount's group) and runs as the first one that can write; only falls back to root if none can. - Named volumes work out of the box.
You can force an identity with the PUID/PGID environment variables.
Configuration
All persistent state (database, logs, settings, feed list) lives in the data
directory: data/ next to the application, in Docker /srv/data (mount it
as a volume). Settings can be changed in the GUI, by editing settings.yaml
there, or with environment variables, which always win over the file.
| YAML key | Environment variable | Default | Description |
|---|---|---|---|
host |
NEWSGETTER_HOST |
0.0.0.0 |
Bind address |
port |
NEWSGETTER_PORT |
8000 |
Listen port |
refresh_interval_hours |
NEWSGETTER_REFRESH_HOURS |
4 |
How often sources are checked |
rejected_retention_days |
NEWSGETTER_REJECTED_DAYS |
7 |
Days to keep rejected articles |
read_retention_days |
NEWSGETTER_READ_DAYS |
21 |
Days to keep read, non-rejected articles |
llm_provider |
NEWSGETTER_LLM_PROVIDER |
chutes |
chutes or opencode |
llm_model |
NEWSGETTER_LLM_MODEL |
— | Model name used for scoring |
llm_api_key |
NEWSGETTER_API_KEY |
— | API key for the LLM provider |
llm_batch_size |
NEWSGETTER_BATCH_SIZE |
15 |
Articles per scoring request |
llm_timeout |
NEWSGETTER_LLM_TIMEOUT |
300 |
Seconds to wait for one LLM request |
llm_no_think |
NEWSGETTER_LLM_NO_THINK |
false |
Append Qwen3's /no_think to prompts |
score_threshold |
NEWSGETTER_SCORE_THRESHOLD |
40 |
Scores below this are rejected by the LLM |
profile_batch_minutes |
NEWSGETTER_PROFILE_BATCH_MINUTES |
45 |
Minutes between batched profile updates |
sources_file |
NEWSGETTER_SOURCES_FILE |
data/sources.yaml |
Path to the feed list |
llm_api_keys (per-provider keys), provider_base_urls and llm_extra_body
(extra request fields, e.g. to disable a reasoning model's thinking mode) are
YAML-only.
The path to the YAML file can be changed with NEWSGETTER_CONFIG.
Adding feeds
The feed list is stored as sources.yaml in the data directory, so it
survives container recreation. Add sources in the GUI or edit the file by
hand:
sources:
- name: Ars Technica
url: https://feeds.arstechnica.com/arstechnica/index
- name: Nature
url: https://www.nature.com/nature.rss
Any RSS/Atom feed works this way.
Entries are deduplicated by URL (tracking parameters are ignored), so re-checking a feed is safe.
YouTube channels
Channels are tracked through their official RSS feed. Use type: youtube
with the channel id (or a playlist id):
sources:
- name: Veritasium
type: youtube
channel_id: UCHnyfMqiRRG1u-2MsSQLbXA
Only the title, the channel name and a short description are stored.
LLM scoring
After every refresh, new articles are scored in batches (score 0–100, a
one-sentence justification, a category) by the configured provider. Scores
at or above score_threshold are accepted; lower ones are rejected by the
LLM. Each article is scored exactly once. A reader interest profile is kept
in the database and updated by the LLM from user actions (read/reject/
restore); signals are batched and applied once per profile_batch_minutes in
a single request. The profile feeds the scoring prompt and can also be edited
manually.
Token usage is recorded per request. Current-month usage per
provider/model/purpose is available at GET /api/usage.
Maintenance
Sources are checked automatically every refresh_interval_hours (changeable
in the GUI — applies immediately). After every cycle new articles are scored
and retention cleanup runs: rejected articles older than
rejected_retention_days and read, non-rejected articles older than
read_retention_days are deleted; unread accepted articles are kept.
Application logs are written to logs/newsgetter.log in the data directory
(with rotation) and can be browsed in the GUI under "Logs". LLM token usage is
shown on the Settings page and at GET /api/usage.
Development
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
hypercorn --reload --bind 127.0.0.1:8000 "app:create_app()"
Project layout
app/ Application code (Quart)
config/ Configuration files
static/ Stylesheets and other static assets
templates/ Jinja2 templates
data/ Runtime data (database) — created automatically