I built a local layer that kills Token Tax–Python lib+Chrome extension+Mac app
Omna is a new tool designed to enhance data accessibility for AI models by providing semantic search and PII masking capabilities. Built on the Polars framework, it allows users to search and mask sensitive data locally without any data egress. The tool aims to bridge the gap between AI and previously unreachable data, making it more usable for various applications.
- ▪Omna enables semantic search and PII masking in a single line of Python code.
- ▪It operates entirely locally, ensuring that no data leaves the user's machine.
- ▪The tool is built on Polars, leveraging Apache Arrow's columnar memory format for efficient data handling.
Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | Hacker News (Newest) |
| Canonical URL | https://omna.dev/ |
| Publication time | Mon, 18 May 2026 01:30:47 +0000 |
| Retrieval time | 2026-05-18T02:03:21.278Z |
| Last seen | 2026-05-18T02:03:21.278Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | V_tEMReTZBfm |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
// 01 — the pitchOmnaSemantic search for Polars.Stop writing regex to scrub patient data.Omna gives Polars semantic search and PII masking — in one line of Python.Local-first. Rust-powered. Zero data egress.Search your DataFrames by meaning, not strings. Mask sensitive columns before they ever reach a model.Local-firstRust kernel0 network callsHIPAA · readyⓘTry it now ↓$pip install omnacopyStar us on GitHub — help us hit 10k ★// 2027 missionWe're building the universal semantic layer between enterprise data and AI. We started with Polars because that's where the fastest-growing data engineering community is.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Hacker News (Newest).