by ContextForge (LX AI)
A context pack is only as trustworthy as the facts in it. If the generator invents your tech stack from a vibe instead of reading the file tree, every agent that reads the file starts from a wrong premise. The safer design puts a deterministic stage first: read the repository, tag what was actually found, and only then let a model draft prose around those facts. ContextForge runs exactly that two-stage pipeline, and the split is what keeps the output honest when the model is unavailable.
The deterministic stage walks the file tree and extracts concrete signals, the package manager, the test command, the lint config, the languages present, and emits them tagged as detected when read from a file or inferred when only guessed from a filename. It never fails and never guesses silently; anything it cannot determine becomes an open question in the output rather than a confident falsehood. These are the facts a model would otherwise fabricate, and they are the ones an agent relies on most.
The model stage drafts the human-readable sections, conventions and pitfalls, using the extracted facts as ground truth. When the upstream model is unavailable, the endpoint returns a 503 with an explicit AI_NOT_CONFIGURED code, and a rule-based draft is returned with a source label rather than a fabricated one. The contract is that a missing key or a failed call never produces text that looks authoritative. A demo mode exists and is marked as such, so no one mistakes it for live output.
Running the model first is tempting because it is fast, but it lets the model fill gaps with plausible errors. Detected facts come from files; inferred facts come from filenames and carry a weaker label; everything else is an open question. When the generated file separates those three, a reviewer can trust the detected block and focus attention on the inferred one. That is the difference between a draft you can sign and a draft you have to re-verify line by line.
No generator, including this one, replaces a human confirming conventions. Every generated context file ends with an open-questions section listing what the tool could not determine, so the handoff to a reviewer is explicit. The pipeline also has no persistence by design, which keeps the tool a privacy feature as much as a convenience. The model assists; the file tree decides what is true.
Useful deterministic signals are boring on purpose. Presence of pnpm-lock.yaml versus package-lock.json tells you the package manager. A pyproject.toml with a pytest config block is stronger than guessing from a tests directory name. Dockerfile base images can hint at runtime, but they should stay labelled inferred unless a matching package manifest confirms them. CI files under .github/workflows often carry the real test command your teammates trust. Directory conventions — apps/, packages/, src/ — help agents navigate, yet they should not invent module boundaries the build system does not enforce.
What stage one must not do is narrate. Narrative belongs in stage two, and only after the fact list is fixed. Mixing narration into extraction is how tools start inventing “this monorepo uses Turborepo exclusively” because a turbo.json once existed in a template. ContextForge keeps extraction output structured: facts with labels, missing items as questions, score breakdown by category (build, test, lint, entry points, stack, conventions, forbidden paths). That structure is what makes a human review feasible in under ten minutes for a typical service repo.
Product teams sometimes paper over outages with a “best effort” paragraph that still reads like live AI. That violates the fleet P0 matrix. ContextForge’s generate endpoint short-circuits on quota before any model call (429), returns 503 when no key is configured (unless demo:true), and 502 when the upstream call fails. Demo mode returns 200 with an explicit banner. Live success carries source: Model-assisted. Rule-based success carries source: Rule-based. Support chat on the site follows the same honesty posture: answers come from the product knowledge base, and compliance topics include a reference-only disclaimer.
For agent security, treat pasted trees like untrusted diffs in a code-review bot. OWASP LLM01 covers prompt injection: user-controlled text must not override system policy. Extraction should ignore instructional comments inside key files when those comments try to rewrite the tool’s job. That does not make ContextForge a penetration test; it is a minimum hygiene bar for any LLM-backed developer tool. See the OWASP Top 10 for LLM Applications for the threat model language your security questionnaire will ask about.
After generation, walk the open-questions list first. Fill build/test/lint commands by running them once locally and pasting the exact strings. Demote any inferred fact you cannot verify. Add forbidden paths for generated code and secrets directories. Commit the pack in the same PR as a small agent task so you can see whether the agent followed the new constraints. If it ignored a rule, tighten the wording or add a CI check — do not add three more adjectives. Re-generate on stack changes so the pack does not become folklore.
This workflow is slower than “ask ChatGPT to write CLAUDE.md from memory,” and that is the point. Memory invents. Trees constrain. ContextForge exists to make the constrained path cheap enough that teams actually use it before Product Hunt demos and before onboarding the next hire onto Cursor or Claude Code.
Posture A: ask a chat model for CLAUDE.md from a short description. Fast, high invention risk, no labels, hard to re-run identically. Posture B: hand-maintain five formats. Accurate when authors care, drifts across tools within weeks. Posture C: deterministic extract plus optional prose, labelled sources, fail-closed statuses. Slower than A, cheaper to keep honest than B. ContextForge implements posture C. Choose A only for throwaway sandboxes. Choose B only if a single owner edits every file. Choose C when multiple agents and multiple humans share the same repo over months.
Byte-stable regeneration matters for CI. If two runs with the same tree produce different detected facts, you cannot gate on drift. Model temperature belongs in stage two only. Stage one should be pure functions over the paste. That is why ContextForge separates analyze and render modules, and why demo mode is labelled instead of silently substituting fiction when keys are missing.
For audits, keep the ruleset version string next to the generated files in your PR description. When security asks what the model saw, answer with the paste policy: tree plus optional key files, no durable archive, fail-closed codes on the API. That answer is shorter and more credible than a glossy security page full of absolute claims.
Suppose the paste shows package.json with scripts.test set to vitest run, a src/index.ts entry, and no eslint config file. Stage one should detect the test command, infer TypeScript from the extension, and leave lint as an open question. Stage two may write a conventions paragraph that mentions vitest, but it must not invent an eslint --fix script. The reviewer then either adds lint to package.json and regenerates, or documents that the repo intentionally skips lint. That single decision is clearer than a model guessing npm run lint because most Node tutorials include it.
Because a model left to infer the stack from a vibe will fabricate it when uncertain. A deterministic first stage extracts real signals from the file tree and tags each fact detected or inferred, so the prose rests on ground truth instead of a guess.
The endpoint returns a 503 with an explicit code rather than inventing content. A rule-based draft can be returned with a source label, and demo mode is marked as such. A missing key never produces text that looks authoritative.
Detected facts are read from a file, such as the test command in package.json. Inferred facts are guessed from a filename and carry a weaker label. Everything else becomes an open question, so a reviewer knows where to focus.
Treat it as a draft. It ends with an open-questions section for what the tool could not determine, and a human should confirm conventions before the file is relied on by an agent.
Try the product: ContextForge
2026-09-16 · primary sources only · no fabricated traffic metrics