Softmax in front of CrossEntropyLoss: 16 other bugs PyTorch won't catch
PyTorch does not catch certain architectural bugs during model design, leading to issues that only appear during or after training. A design-time linter called Neurarch has been developed to detect 17 common structural failure modes in neural networks before training begins. These include incorrect layer ordering, missing components, and inefficient configurations that degrade performance or stability.
- ▪The nn.CrossEntropyLoss function in PyTorch applies log-softmax internally, so adding an explicit Softmax layer causes double application and harms training stability.
- ▪The linter checks for issues such as incorrect normalization order, missing residual connections, and absence of positional encoding in attention layers.
- ▪Rules also flag performance problems like excessive dropout rates and large activation tensors that increase memory usage unnecessarily.
- ▪Some bugs, like placing Dropout before BatchNorm, cancel intended regularization effects and are only detectable through static analysis of the model graph.
- ▪The tool operates on the model's architecture graph before any forward pass, aiming to prevent wasted computation and debugging time.
- ▪Transformer-specific rules catch errors such as incorrect GQA head divisibility and missing auxiliary losses in MoE layers.
Hacker News (Newest) files mainly under programming. We currently carry 5,306 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | Hacker News (Newest) |
| Canonical URL | https://gaox.substack.com/p/how-a-road-network-library-helped |
| Publication time | Sun, 17 May 2026 08:08:10 +0000 |
| Retrieval time | 2026-05-17T08:52:12.990Z |
| Last seen | 2026-05-17T08:52:12.990Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | qrH2ql54DlH5 |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
You can put a Softmax in front of CrossEntropyLoss. PyTorch won’t stop you. Here are 16 other architecture bugs it won’t catch.A walkthrough of the 17-rule design-time linter inside Neurarch: what each rule catches, why it matters, and where static analysis stops being useful for neural networks.Xin GaoMay 17, 2026ShareThe bug that started thisYou can put a Softmax in front of CrossEntropyLoss in PyTorch. The model trains. The loss curve looks fine. You ship it. Accuracy is bad, and you spend the next day finding out why.The bug is that nn.CrossEntropyLoss applies log-softmax internally, so the explicit Softmax causes double-application and degrades training stability. The bug is visible from the architecture diagram in two seconds.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Hacker News (Newest).