Ask HN: What are you using for LLM inference in production?
The author describes their experience using various large language model services, currently relying on Gemini 2.5 Flash Lite for structured-data chatbot tasks. They note that the model is being sunset, leaving them uncertain about future options. They seek recommendations, mentioning Groq and open-source alternatives.
- ▪The user began with OpenAI's GPT‑3 and later experimented with several major AI labs and Workers AI.
- ▪They settled on Gemini 2.5 Flash Lite because it was cheap, sufficiently capable, and fast for chatbot use with structured data.
- ▪The Gemini model is being sunset, prompting the user to look for comparable performance and cost solutions.
- ▪They mention interest in Groq but cite lack of developer access and express a desire for open‑source options.
Hacker News (AI / LLM) files mainly under ai. We currently carry 3,026 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | Ycombinator |
| Canonical URL | https://news.ycombinator.com/item?id=49121047 |
| Publication time | Fri, 31 Jul 2026 09:46:24 +0000 |
| Retrieval time | 2026-07-31T09:52:36.885Z |
| Last seen | 2026-07-31T09:52:36.885Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | yOFLtjwPYY24 · 1 stories |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
I started with OpenAI back in the GPT-3 days, then bounced between the major labs (with a brief interlude with Workers AI).I eventually settled on Gemini 2.5 Flash Lite as my workhorse, using it against structured data/vectors as a chatbot. It was cheap, good enough, and most importantly, fast - but it's now being sunset and I'm not sure where to go to find similar performance.Groq seems promising but developer access hasn't been available for months. I'd love to use open source, but I never found anything comparable (cost/performance/speed).Would love to hear your suggestions.
Excerpt limited to ~120 words for fair-use compliance. The full article is at Ycombinator.