llmXive is a research platform that develops ideas into executable studies, reviews their evidence, and drafts papers for review and publication approval. Specialist LLM agents work through project-specific specifications, plans and tasks. Git records the artifacts, review decisions and recovery history.
The recovery audit found failures in task scope, execution, artifact placement, review protocols and evidence verification. Updating the model alone did not resolve them. The implementation now includes fixes for those observed failures, but a fresh run through the entire authored-paper pipeline is not yet proven.
An earlier isolated totient study autonomously reached research_accepted and
then paper_tasked. Independent checks agreed with all 12 residue-count tables,
24 total-variation rows and 1,000 small-integer identities; its report retained
the observed finite-range exception rather than claiming a general theorem.
A later paper-task failure routed that run to the wrong research stage, so it
was stopped and preserved. Its scientific artifacts and state were not manually
edited to obtain acceptance. See the recorded evidence.
That result establishes useful research-stage progress, not an accepted paper or recovery of the production backlog. A subsequent fresh full-pipeline run remains the acceptance test. Reviewed external preprints, green CI, model probes and successful individual repairs to the platform are distinct evidence. Follow #1139 for current throughput and full-pipeline results, and #1242 for autonomous repair: no accepted useful autonomous platform repair has yet been demonstrated.
The lifecycle graph schedules agents from the registry, currently 54 entries. Each project owns its Spec Kit scaffolds, research code, data, paper, reviews and revision history. The registry and central backend router define model behavior; state files record progress. The dashboard derives its view from these records.
| Phase | Main path | Required evidence |
|---|---|---|
| Research design | Idea → expansion and question validation → specification → clarification → plan → tasks and analysis | A scoped, testable question; complete design artifacts; executable tasks; stage review |
| Research execution | in_progress → research_complete → research_review → research_accepted |
Actual code/data outputs, current execution and test evidence, verified claims/references, and unanimous research-panel acceptance |
| Paper development | paper_drafting_init → paper specification, clarification, plan, tasks and analysis → paper_in_progress → paper_complete |
A manuscript and figures grounded in the accepted research, runnable build instructions, and compiled output |
| Publication | paper_review → paper_accepted → awaiting_publication_signoff → posted |
Current paper-panel acceptance, explicit maintainer approval, and successful publication |
The complete state vocabulary is defined in types.py. Failures do not count as advancement: they retain diagnostics and either retry, reopen the responsible task, or return the project to the appropriate design stage. Paper task-format, planning and writing failures retain the paper track instead of corrupting accepted research state; substantive scientific findings can still require a return to research or the backlog. Successful analysis can be reused only when its recorded artifacts and relevant review inputs still match.
Reviewable stages use one convergence engine:
identify actionable concerns, revise against those concerns, then re-review the
same concern set. The configured round cap bounds a review attempt; it does not
guarantee acceptance. Unresolved substantive concerns cause a documented
kickback. Revision work lives under
projects/<PROJ-ID>/.specify/auto-revisions/round-<N>/; prior rounds and archives
retain provenance.
Research and paper review require unanimous panel acceptance with no open concerns. Document-authoring stages may advance with writing-only residue; requirement, methodology or scientific defects still block them. Human and explicitly labeled simulated-personality comments inform the relevant reviewer but do not directly substitute for a specialist verdict. Publication has a separate mandatory human sign-off.
Independent task verification examines the actual referenced artifacts and their canonical project paths. Execution evidence binds to the run-book, source and output bytes; changed or missing evidence invalidates approval. Signed empirical result receipts authenticate captured bytes and provenance, not scientific correctness. Citation and claim checks, independent review and final execution checks remain separate gates.
After paper acceptance, the publication workflow assembles the review trail and requests maintainer approval through a GitHub issue or the publication CLI. Maintainer rejection returns feedback to revision. Approved publication uses the Zenodo publisher and records the DOI and publication metadata. This is implemented behavior, not evidence that the current fresh canary has completed publication.
Submitted or imported papers follow a separate paper_ingested →
reviewed_preprint path. The original scientific work is preserved, review
artifacts are attached, and a separate brainstormed project may explore a
follow-up question. Reviewing someone else's preprint is not completion of an
llmXive-authored study.
The personality workflow currently runs every two hours. Its persona prompts draw on public writings; contributions are labeled “(simulated)” and are advisory. The attribution audit checks the labeling invariant.
Dartmouth Chat is the production model service. The central default and all
LLM-agent registry defaults use GLM-5.3 (zai-org.glm-5.3). For transient
failures, the router tries free same-backend peers in order:
GPT-OSS 120B (openai.gpt-oss-120b), then Gemma 4 31B
(google.gemma-4-31B-it). Catalog availability and returned-model identity are
checked separately from configuration. A fallback response is recorded as that
model; primary-model acceptance tests do not let a peer impersonate the primary.
The router also supports backend fallback where configured. Local transformers requires installed dependencies, suitable hardware and a compatible available model; it is not a guarantee that a large Dartmouth model can run on a laptop. Qwen remains a compatibility route for callers explicitly requesting it, not the production default. Specialized vision, personality and manual utilities still have explicit model choices; registry defaults do not describe every call. Those remaining choices are tracked in #1285.
The existing paid Dartmouth fallback is off by default. It requires
LLMXIVE_PAID_OPT_IN=1 and headroom under the
credit-budget guard; an unavailable or invalid
balance check refuses paid calls. The deployed
advance workers explicitly enable that opt-in
with LLMXIVE_PAID_BUDGET_FRACTION=0.9, so production can use the guarded fallback
within Dartmouth's credit budget after free peers fail. This is separate from
the HF pilot. Included credits and endpoint availability are service/account
properties, not a promise of unlimited free inference.
Research designs should fit available compute. The existing
Kaggle offload adapter can submit eligible
GPU-bound execution when configured with KAGGLE_API_TOKEN, then poll and collect
outputs. Missing credentials or unavailable capacity do not establish successful
execution. The optional HF pilot below is a separate, explicitly invoked path.
The contextlab Academia pilot has private versioned artifact storage and a
private working bucket in an llmxive resource group. Jobs and Inference
Providers share one $20 monthly compute cap; paid Spaces and Inference Endpoints
have $0 limits. This is a pooled allocation, not ten independent wallets or an
automatic production fallback. Provider enforcement can lag the cap.
The pilot guide records successful private-file upload/download verification, a bounded CPU computation and one alternate-model call. A recovery-audit copy was also restored byte-for-byte with a pinned storage receipt. The working bucket's data round trip remains untested. Automatic research-artifact migration and production compute offloading remain future integration work.
The helper reads HF_TOKEN from a CI secret or the local macOS Keychain; it does
not read plaintext HF caches or repository .env files. The pilot currently uses
Keychain, with no recurring HF workflow or pilot HF CI secret enabled. Keep all
tokens out of Git, command arguments and logs. Use the explicit commands in the
pilot guide only when a bounded storage, compute or model experiment is needed.
Generated research runs in a project-specific Python environment with a project-local working directory and runtime home. Its environment does not inherit orchestrator credentials. Python and supported shell commands have time limits and captured output; failed execution feeds the next correction attempt. These process safeguards are not a complete hostile-code filesystem sandbox: research jobs may access the network to obtain public data. Run studies in disposable workers/checkouts with appropriate host permissions.
Implementation and planning writers check project boundaries and symlink escapes.
Planning refuses ambiguous file aliases before writing and restores prior bytes
when deterministic guards reject a revision. Implementation refuses new
case/Unicode-equivalent filename collisions.
Default scheduled commits permit project paths, state/ and web/data/;
platform repairs use a separate reviewed path. The trusted Pages workflow selects
a distinct profile that stages only docs/.
agents/ Registry and prompts
src/llmxive/ Pipeline, agents, backends, execution, review and repair
projects/<PROJ-ID>/
idea/ Research seed and expansion
specs/ Active research specification, plan, tasks and contracts
code/, src/, scripts/ Project implementations (layout depends on the study)
data/, figures/, results/ Project inputs and generated evidence
paper/ Paper scaffold, manuscript, figures and build outputs
reviews/ Review records and advisory feedback
.specify/ Project controls, revision rounds, caches and archives
state/ Project state, run logs and verification records
specs/ Platform-development specifications
web/ Dashboard source and derived data
docs/ Deployed dashboard mirror; do not hand-edit
scripts/hf/ Explicit bounded HF pilot and receipts
eval/promptfoo/ Repeated prompt-regression evaluations
tests/ Unit, contract, integration and live checks
.github/workflows/ Scheduled work, checks, repair candidates and publication
The root-file recovery audit preserved 2,666 misplaced files with their original Git bytes and modes. Known owners received project-local recovered downloads; uncertain owners remain in provenance-addressed recovery storage. Recovery does not validate a dataset for scientific use. Remaining historical filename collisions are inventoried, including different scientific files that must not be silently overwritten.
Run the fast tracked-layout check with:
python src/llmxive/checks/repository_layout.pyUse Python 3.11 or newer in a dedicated environment. Paper builds require a TeX toolchain; PDF audits require Poppler. Follow the relevant workflow's dependencies for the operation you are running.
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
python -m llmxive auth set # optional interactive local Dartmouth setup
python -m llmxive preflight
python -m llmxive brainstorm -n 5
python -m llmxive run --project PROJ-ID --max-tasks 1PROJ-ID is an existing project's full identifier. Omitting --project lets the
scheduler choose eligible work. These commands make real model calls and persist
project changes; use an isolated checkout for experiments. See
python -m llmxive --help for submission processing, direct agent invocation and
project operations. Maintainers can record a publication decision with:
python -m llmxive project publish-approve PROJ-ID \
--who 'Maintainer Name' --what 'reviewed paper meets standards'GitHub Actions supplies service credentials through repository secrets. For
Dartmouth, the CLI also supports a permission-checked local credential file via
auth set; HF pilot credentials follow the stricter Keychain/CI-only path above.
Research execution needs a persistent LLMXIVE_RECEIPT_KEY in the orchestrator
and workers verifying its results. The supported local credential field is
llmxive_receipt_key in ~/.config/llmxive/credentials.toml with mode 0600.
Changing the key invalidates old receipts. Keep signing keys outside the repository
and generated research subprocesses. Zenodo publication requires its own token;
sandbox publication tests use a separate sandbox token.
The main pipeline is scheduled every three hours, with additional advance workers and stage-specific workflows. Submission intake is hourly; publication sign-off polling is every two hours. Workflow files are the authoritative schedules. A successful worker exit can mean no eligible progress; inspect stage transitions and artifacts.
The repair workflow collects concrete runtime errors or actionable issues and proposes a bounded platform change. It requires a regression that exercises an existing production caller and fails on the baseline, passing candidate and related tests, fixed preservation tests, and review by a different model. Invalid syntax, invented-helper import failures and deletion of prior user content do not count as successful repairs.
Generated repair tests run in separate resource-limited Docker containers with network disabled, no credentials or Docker socket, read-only inputs and temporary working storage. The candidate cannot replace existing tests. A separate publication job verifies the evidence digest and unchanged target-file baseline before opening a reviewable PR; the candidate generator does not directly deploy its own patch. See the repair preservation evidence.
These safeguards have rejected bad candidates. They do not yet demonstrate an accepted, useful autonomous repair. Follow #1242 for actual candidate outcomes rather than treating a scheduled run as improvement.
pytest tests/contract
pytest tests/unit -m 'not slow'
pytest tests/integration
LLMXIVE_REAL_TESTS=1 pytest tests/real_call -m 'not slow and not external_references'Live suites require the relevant credentials and services. The PR workflow runs offline and selected live jobs concurrently. Dartmouth checks cover the actual configured primary and runtime behavior. External reference/dataset availability is checked separately when changed dependencies warrant it; known documentation and orchestration changes do not need unrelated registrar uptime. Unknown inputs remain conservatively selected. The selector and dependency guards keep this policy reviewable.
The nightly suite retains slow checks, free-peer coverage, external services and full-pipeline acceptance tests, with failure artifacts preserved. Prompt Eval runs two reviewer fixtures three times, including rejection of a deliberately flawed idea, using the configured primary and rejecting peer-model substitution. Repository audits check layout, Spec Kit artifacts, PDFs, personality attribution and feedback behavior. Sparse checkouts reduce data transfer while preserving each audit's required input corpus.
Tests establish specific contracts. Full scientific acceptance additionally requires an observed run through research execution, paper production and review, with useful artifacts that survive independent inspection.
The dashboard is a static view of canonical
project state, built by web_data.py. Edit web/;
Deploy Pages regenerates and publishes docs/.
It displays authored projects separately from reviewed external preprints.
Use the dashboard to browse artifacts and review trails, submit an idea or paper, and provide feedback. Contributions can also start with a GitHub issue or a pull request. For development, follow an entry in agents/registry.yaml to its prompt and implementation; platform specifications live in specs/. Recurring failure evidence belongs in the existing tracking issue when possible.
See LICENSE. Maintained by the Contextual Dynamics Laboratory at Dartmouth College.