Search public sources. Analyze the data locally. Export the evidence.
The Windows app collects public records and datasets, extracts text, tables, and repeated structured fields into rows, and analyzes them on your computer. The browser demonstration below runs five prepared questions with the same analysis engine.
Five prepared questions · runs locally on 14,999 rows
Choose or type one of the five sample questions.
Engine booting in your browser…
- loading columns…
The panel will show the computed group statistics when the validated engine finishes booting.
Analyze a CSV on this device.
Move through Data, Question, Methodology, Results, and Export. Your session stays in this browser.
Restored your last session (local only).
Data
Nothing is uploaded. If you choose to remember it, the file stays in this browser's storage until you clear it. Remembered files expire after 30 days. CSV rows stay on this device while the workbench is open.
The analysis engine is loading the bundled sample.
| Column | Detected role | Missing | Example values | Role override |
|---|
Role overrides guide engine recommendations, diagnostics, and eligible column controls.
Question
Proposed analysis
Runnable suggestions
Review the full method catalog
Methodology
Use Ctrl or Command to select more than one predictor.
The engine receives this level when the selected method accepts it.
No engine counterpart is available for this method.
Methodology document preview
Results
Export
Run an analysis before building its reproducibility package.
Package contents
- Prepared analysis data and SHA-256 hashes
- Methodology, caveats, and environment details
- Generated Python and illustrative R code
- Machine-readable results, assumptions, and chart SVG files
- The supplied collection receipt, or a missing-receipt note
Do employees who left report lower satisfaction?
This fixed 14,999-row HR example records one Welch t-test. It stays fixed while the live workbench above changes method. Its methodology, data, runnable check, result, and machine receipt are available together below.
Dataset attribution and verification limits appear in the methodology below.
Employees who left reported mean satisfaction of 0.440, compared with 0.667 among employees who stayed.
t(5167.03) = -46.64, p < .001, d = -0.99
This observational sample shows an association. It does not establish why employees left.
- Data
- 14,999 rows, 10 columns
- Groups
- 3,571 left, 11,428 stayed
- Method
- Welch independent t-test
- 95% interval
- [-0.236, -0.217]
- Engine
- Validated local analysis core
Methodology and assumptions two-sided Welch t-test
Question and variables
The outcome is satisfaction_level, a numeric score from 0 to 1. The grouping variable is left, coded 1 for employees recorded as having left and 0 for employees recorded as having stayed.
Dataset source
This case uses a 14,999-row HR Analytics sample formerly distributed through Kaggle. The original listing is no longer available. Current package documentation records the Kaggle origin and schema. The receipt records the unavailable listing, local file hash, column names, and verification limits. The original publisher and dataset license have not been independently verified.
Why this method
The analysis compares the two group means with a two-sided Welch independent-samples t-test. Welch's test allows unequal variances and unequal group sizes. The reported difference is left minus stayed.
What is checked
The runnable check verifies the CSV's SHA-256 hash, recomputes both group means, t, degrees of freedom, p, Cohen's d, mean difference, 95% interval, and APA line, then fails if a checked value differs from the receipt.
Limits
Rows are treated as independent and satisfaction as numeric. Dependence, repeated employees, selection bias, measurement error, or undocumented data generation could change the estimate. Interpret this comparison only within the observed sample.
From question to export in four steps.
Each step stands on its own. Bring data and stop, collect and export, or run the whole path.
-
Bring or collect data
Load a spreadsheet, search public sources, or paste a page. The collector preserves the source URL, retrieval time, and method with each record.
-
Ask in plain language
Type your question. The assistant maps it to a validated method and asks what you mean when a choice changes the answer.
-
The engine computes
A deterministic engine runs the statistics on your machine and returns each number with the group statistics and its APA line.
-
Export the analysis
Export the data, method, parameters, result, and runnable files another person can use to repeat the analysis.
Collect records, run statistics, and export the evidence.
Pick a capability to inspect its inputs, methods, and output record.
Analyze
Ask in plain language or set the variables yourself. The engine includes reference-validated t-tests, ANOVA, correlation, regression, and other methods. Every result states its trust tier and carries the statistics behind it.
The guided builder writes real SQL for a non-programmer. The expert picks the test, the variables, and the parameters directly.
Sample data
hr_analytics.csv| loading dataset… |
Collect
One research run can search Reddit, Hacker News, GitHub, Polymarket, and the general web. YouTube joins when yt-dlp or a ScrapeCreators key is available. Optional keys add X, TikTok, Instagram, and other provider-backed sources.
Full-depth web search requests up to 20 results per query. Page extraction defaults to 20 URLs, prioritizes distinct domains, and can be raised to 50. Direct-page collection reads HTML tables, repeated page records, readable text, JSON-LD, and microdata into structured rows. Requests are paced per site, check robots.txt by default, honor Retry-After, and stop that site when verification or access controls appear.
Sources
public and BYO key- Zero configurationReddit, HN, GitHub, Polymarket
- General websearch results with source URLs
- Direct pagestables, text, structured data
- Open datapublic datasets and APIs
- Provider sourcesyour API key, your limits
Prove
Every result carries an honest label. Validated methods are checked against reference tools and covered by regression fixtures. Experimental methods carry a separate label.
When a request falls outside both tiers, the optional code-authoring path shows editable Python before an explicit run. It is not an OS sandbox, runs with the user's Windows permissions, and remains labeled AI-authored and not reference-validated.
Trust tiers
a label per number- Validatedchecked against reference tools
- Experimentalimplemented, labeled
- AI-authoredshown, then run locally
Trace each result from source to output.
Collabysis records the source, method, parameters, and output so another person can inspect the path from row to result.
Checked against reference tools
Validated methods are tested against frozen reference values and, where available, SciPy or statsmodels oracles.
A pack that re-runs to the same figures
Every validated run can export a reproducibility pack containing the data, method, parameters, checks, and runnable files needed to recompute the checked figures.
Your data stays on this computer
Imported data and analysis stay local. If AI assist is enabled, it sends the question and column names and roles to the provider the user selects; data rows are not sent by that feature.
What the evidence pack records
A trace from collection to export.
| Stage | Recorded evidence | Why it matters |
|---|---|---|
| CollectionPublic source retrieval | Source URL, timestamp, method | Shows where each row came from |
| PreparationChanges before analysis | Filters, column roles, exclusions | Makes cleaning decisions visible |
| AnalysisThe calculation | Method, variables, parameters, backend | Lets a reviewer repeat the test |
| ExportFiles for review | Dataset, report, code, manifest, receipt | Keeps the result attached to its evidence |
Install Collabysis on your PC.
The desktop app adds multi-source collection, direct-page extraction, durable private projects, and Windows workflows. It runs on your computer.
Version 3.8.3 · 67.7 MiB. The early-access installer is unsigned, so Windows may show an Unknown publisher warning. Verify the published SHA-256 checksum before running it.
Version 3.8.3 is licensed under the Collabysis Software License Agreement, version 1.0. LICENSE.txt and THIRD_PARTY_NOTICES.txt are inside the ZIP. Analyses and collections run with a 30-day trial key or a paid key valid for one year from purchase. Opening existing files, exporting existing results, viewing settings, and checking for updates work without a key. This build stops running analyses and collections on 2026-12-15; this page always offers the current build. Release packaging rejects .env files, private-key files, and detected credentials. Data collected from third-party sites keeps its original rights and terms.
Product and license details.
Does my data leave my computer?
Analysis stays on this computer. Collection and optional AI assist contact the services you select. See privacy and data flow for the fields each service receives.
Is Collabysis free?
The stock desktop build includes 30 metered analysis actions for the lifetime of one installation. The counter is a local courtesy gate for this allowance. It is not tamper-resistant licensing. It does not use or store an IP address. A valid offline license bypasses it. Opening existing work and downloading existing files use no action.
Do I need to know statistics?
The workbench explains method inputs, assumptions, effect estimates, and trust labels. You can ask a supported plain-language question or select the variables and method directly.
Does Collabysis scrape complete web pages?
Yes when you use the direct-page collector. It extracts readable text, HTML tables, JSON-LD, and microdata. Optional bounded discovery can propose page links, feeds, and sitemaps. Direct requests are paced per site and respect robots.txt by default. Transient failures use bounded retries. A CAPTCHA, verification page, login wall, or persistent access block stops that site. Collabysis does not evade those controls.
Can it collect millions of rows?
A permitted source API, sufficient quota, and available storage can support large datasets. Volume alone says nothing about relevance, independence, coverage, or quality. The custom-link collector uses bounded discovery and keeps one run in memory. This release has no durable checkpointed crawler.
What can it analyze?
The desktop app includes reference-validated and experimental statistical methods, descriptive summaries, queries, and labeled models. Requests outside those catalogs can use an optional AI-authored path that shows editable code before an explicit run and labels the result separately.
Which platforms does it run on?
The downloadable early-access build supports 64-bit Windows 10 and Windows 11. The browser workbench runs in current desktop browsers. A Mac installer is unavailable.
Can someone use it without installing the desktop app?
Yes. The browser workbench imports CSV files, runs supported local analyses, and exports reproducibility packages. Public-source collection remains in the desktop app.
What license applies, and what is in the ZIP?
Version 3.8.3 is licensed under the Collabysis Software License Agreement, version 1.0; LICENSE.txt is inside the ZIP. The hosted browser workbench is free and needs no key. Analysis and collection in the downloaded Windows app need a trial or paid key. Results, charts, generated code, and reproducibility packages created with the software belong to the person who created them and may be published without restriction. Third-party components keep their own licenses. See the third-party notices.