Telegram
MAX.
Posts.
Models.
Filters.
Posts from the MAX folder.
Classification, filters, source attribution and publishing.
Production overview · 3 October 2026 · counters at capture time
What happens
to a post
Collection
Telethon receives new messages from the MAX folder and saves the original.
Model
TF-IDF computes scores. MiniLM can be run separately for comparison.
Filters
Conditions check scores and length. A match assigns a dictionary label.
Source
A source link is added. Attachments and freshness are checked.
MAX
The sender selects a channel by label and saves the delivery result.
The board has five stages: Received, Sorted, Filtered, Source attribution and Ready. Unmatched posts stay in Sorted. Successful deliveries appear in a separate history.
The original text is sent with attachments. The original URL sits behind the word Source. This route has no rewrite step or extra extracted-links block.
Project screens
Captured from the running service. The gallery shows public source material and processing state. Click to open the full-size image.
Channels
in MAX
One post can fit multiple topics. Its label selects the route.
Beginner and humor selection is still experimental. Misclassifications are possible; I tune the filters manually.
Example published post ↗What runs where
Computer
Corpus, Codex annotation, training and comparison of TF-IDF / MiniLM. Raw data and secrets stay outside Git.
VPS
Telethon · FastAPI / Jinja · PostgreSQL · SQLAlchemy · RapidOCR / ONNX · Docker / Nginx.
MAX
A separate sender uploads media, publishes, saves IDs and verifies delivery. The working interface requires a password.
Code-map
Pinned source snapshot · 5f50077b
app/live_collector.pyReceive and reconcile new messagesapp/pipeline_coordinator.pyStage transitions and job creationapp/taxonomy/jobs.pyModel queue and input sourceapp/taxonomy/inference.pyClassifier scoresapp/ocr/worker.pySeparate OCR workerapp/content/selection_rules.pyFilter condition treeapp/content/selection_filters.pyMatches, labels and evidenceapp/publication/worker.pySend and verify deliveryWhat still needs work
- Corpus labels were assigned by Codex. Model measurements describe agreement with those labels, not independent human accuracy.
- The beginner and humor routes are experimental. Beginners: 57/65 test matches; OCR humor: 92/104 validation matches. The agreed quality gates were not passed.
- OCR does not understand an image without text, listen to audio or inspect the whole video. Reading errors need review.
- Media over 100 MB and unavailable album parts block automatic readiness. The interface contains stopped cards.
- Screenshot counters and timings are a state snapshot, not a performance guarantee.