re:publisher
EN/RU

Telegram
MAX.

Posts.
Models.
Filters.

My project for topic feeds.

Posts from the MAX folder.
Classification, filters, source attribution and publishing.

Production overview · 3 October 2026 · counters at capture time

What happens
to a post

01

Collection

Telethon receives new messages from the MAX folder and saves the original.

02

Model

TF-IDF computes scores. MiniLM can be run separately for comparison.

03

Filters

Conditions check scores and length. A match assigns a dictionary label.

04

Source

A source link is added. Attachments and freshness are checked.

05

MAX

The sender selects a channel by label and saves the delivery result.

The board has five stages: Received, Sorted, Filtered, Source attribution and Ready. Unmatched posts stay in Sorted. Successful deliveries appear in a separate history.

The original text is sent with attachments. The original URL sits behind the word Source. This route has no rewrite step or extra extracted-links block.

OCR and post text are scored separatelyText from images and video previews. The caption is not mixed into the OCR input.

Project screens

Captured from the running service. The gallery shows public source material and processing state. Click to open the full-size image.

01Overview: stages, service heartbeats, model queues and media errors. Production capture, 3 October 2026.
Full size ↗
02MAX publishing is enabled. Delivery counters and routes are shown at the time of capture.
Full size ↗
03A public Telegram post: original attachment, separate processing and a saved MAX delivery.
Full size ↗
04An empty caption and text on an image: expanded OCR and separate TF-IDF scores. This example has no MiniLM run.
Full size ↗
05The active Humor · OCR filter: caption length, OCR score input and a user-defined label.
Full size ↗
06Five board stages. Filtered by NVIDIA and the Tools label; a public source post is visible.
Full size ↗
07A public post card: saved TF-IDF runs on the original text and deliveries to multiple channels.
Full size ↗
08The MAX result: original text, attachment and a small Source link. Captured from a published channel.
Full size ↗

What runs where

01

Computer

Corpus, Codex annotation, training and comparison of TF-IDF / MiniLM. Raw data and secrets stay outside Git.

02

VPS

Telethon · FastAPI / Jinja · PostgreSQL · SQLAlchemy · RapidOCR / ONNX · Docker / Nginx.

03

MAX

A separate sender uploads media, publishes, saves IDs and verifies delivery. The working interface requires a password.

Code-map

Pinned source snapshot · 5f50077b

What still needs work
  • Corpus labels were assigned by Codex. Model measurements describe agreement with those labels, not independent human accuracy.
  • The beginner and humor routes are experimental. Beginners: 57/65 test matches; OCR humor: 92/104 validation matches. The agreed quality gates were not passed.
  • OCR does not understand an image without text, listen to audio or inspect the whole video. Reading errors need review.
  • Media over 100 MB and unavailable album parts block automatic readiness. The interface contains stopped cards.
  • Screenshot counters and timings are a state snapshot, not a performance guarantee.