WeSearch

Softmax in front of CrossEntropyLoss: 16 other bugs PyTorch won't catch

Xin Gao· ·7 min read · 0 reactions · 0 comments · 26 views
#machine learning#pytorch#neural networks#software linter#deep learning
Softmax in front of CrossEntropyLoss: 16 other bugs PyTorch won't catch
TL;DR · WeSearch summary

PyTorch does not catch certain architectural bugs during model design, leading to issues that only appear during or after training. A design-time linter called Neurarch has been developed to detect 17 common structural failure modes in neural networks before training begins. These include incorrect layer ordering, missing components, and inefficient configurations that degrade performance or stability.

Key facts
About this source

Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.

Original article
Hacker News (Newest) · Xin Gao
Read full at Hacker News (Newest) →

Story provenance

Source · retrieval · rights · ranking — open for full record
inspect →

Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.

Record

Original publisherHacker News (Newest)
Canonical URLhttps://gaox.substack.com/p/how-a-road-network-library-helped
Publication timeSun, 17 May 2026 08:08:10 +0000
Retrieval time2026-05-17T08:52:12.990Z
Last seen2026-05-17T08:52:12.990Z
Headline sourcePublisher (no WeSearch rewrite)
Excerpt sourcepublisher body
Excerpt methodFirst ~120 words (~800 chars) of extracted publisher body, fair-use limited.
SummaryWeSearch · cerebras-chat (WeSearch summarizer)
Summary source textcontentText
Citation coverageSummary is a WeSearch-generated derivative; primary citation is the original publisher URL.
ClusterqrH2ql54DlH5
Cluster logicGrouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison.
Ranking reasonStory pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking.
Publisher visitYes — open original
Substitutes article?No — link-out required for full text

Rights status (four layers)

Publisher-declared
No publisher-confirmed rights record for this source yet.
Machine-readable
No source-specific machine-readable restriction detected beyond the public feed.
WeSearch interpretation
WeSearch declared handling (basis: Derived from the published RSS/Atom feed). This is WeSearch policy, not a legal grant on the publisher's behalf.
Unknown
Retrieval and training permissions are not asserted unless the publisher confirms them.

WeSearch handling by dimension

Indexing May the item be indexed (stored, ranked, made findable)? Allowed
Snippet May a short excerpt of the publisher's text be shown? Allowed
AI summary May WeSearch generate its own short summary of the article? Limited
Retrieval / RAG May the content be exposed for third-party retrieval-augmented generation? Not asserted
Model training May the content be used to train AI models? Not asserted
Commercial reuse May the content be reused commercially? Not permitted

Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.

Opening excerpt (first ~120 words) tap to expand

You can put a Softmax in front of CrossEntropyLoss. PyTorch won’t stop you. Here are 16 other architecture bugs it won’t catch.A walkthrough of the 17-rule design-time linter inside Neurarch: what each rule catches, why it matters, and where static analysis stops being useful for neural networks.Xin GaoMay 17, 2026ShareThe bug that started thisYou can put a Softmax in front of CrossEntropyLoss in PyTorch. The model trains. The loss curve looks fine. You ship it. Accuracy is bad, and you spend the next day finding out why.The bug is that nn.CrossEntropyLoss applies log-softmax internally, so the explicit Softmax causes double-application and degrades training stability. The bug is visible from the architecture diagram in two seconds.

Excerpt limited to ~120 words for fair-use compliance. The full article is at Hacker News (Newest).

Anonymous · no account needed
Share 𝕏 Facebook Reddit LinkedIn Threads WhatsApp Bluesky Mastodon Email

Discussion

0 comments