OrcusBack to the workspace

Terms of use

These terms cover the Orcus formulation service. Using the service means accepting them. They describe what the service does, what has been measured about it, what it does not claim, and what happens to the information you put into it.

1. What the service is

You describe an experiment and list the materials on your shelf, and the service plans the round. It is a planning tool for experiments you then run yourself: it is not a laboratory service, it does not supply materials, and it performs no measurement of any kind.

Orcus chooses which formulations to make next. It builds the recipes your materials allow, chooses the batch worth making next, and gives it to you with the masses and volumes for the bench. The code that does this for you is the code that was measured: replayed against three published screens it beat picking at random on all three, with a margin of +4.90 enrichment at the largest. Those screens are retrospective, so only formulations somebody had already made could be recommended. The part that invents new recipes cannot be tested that way and was not tested there.

The first batch is not a ranking. With no measurements there is nothing for a model to learn from, so the first batch is a designed screen that covers the possibilities as widely as it can. The ranking starts in round two, when your own numbers come back. This is a property of the method and not a setting. Round one is model free by construction: with nothing revealed, every model predicts the same value and the batch is decided by the design alone.

It is built for choosing ratios of materials you already have, which is the problem the prior can speak to. Of 54 published screens large enough to work with, 3 vary composition independently of the lipid, and those are the ones this was developed and measured on. Choosing a new ionizable lipid is a different problem and this does not address it. Three screens is a small base, and two of the three are the development data for most of the numbers below.

Every recipe is checked against your own stocks and scale before you see it, and what fails is reported rather than quietly dropped. On the reference inventory at the default scale, 39 of 48 recipes were makeable, and the rest were reported with the reason. The check is about arithmetic and glassware. It says nothing about whether a recipe works.

2. What you are responsible for

You remain responsible for the science and the safety of anything you make. The service works from the materials, limits and process values you enter, and cannot know whether a value is right. It checks arithmetic, masses, volumes and your own declared limits, and reports what it cannot check. A batch it returns has not been reviewed by a person, and nothing in it removes the need for your own judgement, your own risk assessment, or the approvals your institution requires.

  • Check every recipe against your own protocols before making it.
  • Leave a limit blank when your lab does not have one, rather than entering a typical value: a guessed limit silently narrows the search.
  • Recipes you edit yourself are marked as yours, and the score no longer describes them.

3. What has been measured

The statements in this section are generated from the runs that produced them, so they cannot drift from the numbers behind them. The results that went against the method are here with the ones that went for it.

Against picking formulations at random, the selector finds better ones. Over 36 slices of published screens it reached an enrichment of 1.583 where random reached 0.959, winning 32/36 slices, with a corrected p of 7.377e-06. Enrichment of 1.0 is what random luck gives, so 1.583 means that many times better than luck. This is data the project had already seen, and no pre-registration changes that. The 36 slices come from two publications, so there are two independent units, not 36.

Against a mixture design from 1958, which is the standard way a formulation scientist would lay out the same experiment, no difference was established. The selector scored 1.583 and the classical design 1.589, and the selector took 14/36 slices. The corrected p is 0.7455, which is far from the bar fixed before the run. It is not behind either. No difference established means the comparison could not tell them apart, and there is no confidence interval for the margin anywhere in the record. If you want the simplest thing that works, the classical design is not worse on this evidence.

On the one published screen where both the lipid and the ratios vary, it beat reproductions of the published machine learning methods. Enrichment 2.280 against 1.313 for the LANTERN style network and 1.092 for the AGILE style model, over 30 seeds, every corrected p at or below 0.0008. One screen is one unit. These are reproductions of the published architectures and not the authors' own code, trained from scratch on what the loop reveals, which is the setting a lab is in but is harder on a method built to be pretrained. Those methods were built for ranking new lipids, not for choosing ratios, so a loss here does not refute those papers.

On the set held back from the work, it is not in front. Orcus found 5.12 distinct good formulations per publication against 5.27 for gradient boosting on molecular fingerprints. Against the rival named in advance the win count was 4/5, which the plan had already agreed to call weak support. With five publications the smallest p this design can return is 0.0625, so nothing here can reach the usual threshold, and that was written down before the run. The set was also not untouched: 40 of the 42 publications in the frame appear somewhere in the project's own results.

Reading the literature has never improved a ranking here. Three ways of joining a paper to a formulation were tried, by ratio, by lipid name and by chemical structure, and none of them helped. The best of them scored 1.196 against 1.454 for the same model with no literature at all, over 38 units from 32 publications, taking 16 of 38 units at a corrected p of 1. The point estimate is the wrong way round: carrying the literature cost enrichment rather than adding it. Not significant in either direction, so the honest statement is that it buys nothing. Saying it has no effect would be generous, because the best estimate available is that it costs. Papers are still retrieved and shown to you, and they are never an input to the ranking.

One reason for that is measured rather than guessed. On a random sample of 20 sentences the first reader called results, its direction was wrong on 13. Papers are now read with the claim checked against the paper's own text, and the readings you see are labelled by which reader produced them. Correct reading did not rescue the ranking. Papers barely vary the one thing the prior was keyed on, so even correct readings could not tell candidates apart. That is why this work ends in a panel you can read rather than in the model.

Molecular structure was tried as an input and made the batch worse. On lipids the model had never seen, adding fingerprints moved enrichment by -4.32 at a p of 0.0006, while it moved rank correlation the other way by +0.143. Better prediction overall and a worse top of the list is a real result, and the top of the list is the whole job. One screen can answer this, because it is the only one whose ionizable lipid varies. Structure was dropped for discovery on this evidence, not declared useless in general.

4. What is not claimed

The service makes no claim about how well a formulation will perform. The score orders candidates so that a round of experiments spends its wells on the more promising ones, and that is the whole of it.

  • No formulation from this system has ever been made in a lab. There are no measurements from a bench anywhere in this record, and nothing here has been confirmed by an experiment.
  • The score does not predict performance. It ranks candidates against each other for the purpose of choosing which to make next. Particle size, uniformity, encapsulation, potency and toxicity are not predicted, in any column.
  • Nothing here has been checked by anyone outside the project. No external reviewer, no independent replication, and no prospective test on data that did not exist when the method was built.
  • It is not the best method available. It matches a classical mixture design on development data and sits behind fingerprint gradient boosting on the held back set, 5.12 against 5.27 distinct formulations per publication.
  • It does not judge whether a material can be synthesised, whether a formulation is patentable, or whether anyone has made it before. Already tested means absent from the measurements you supplied, and nothing more.

Each of those has a limit worth reading with it. In short: every number in section 3 comes from replaying the method against published screens, not from a bench; the endpoints a lab cares about are recorded too rarely in the data behind the prior for anything to learn them; and nothing here has been checked by anyone outside the project.

Every number in this document is a replay against published screens. Whether a scientist accepts a recipe, whether it forms particles at the intended size and encapsulation, and whether it delivers are all open, and a lab is the only thing that can close them.

5. Your data, and what is learned from it

What you enter stays yours. Your brief, your materials and your measurements are used to run the round you asked for, and for nothing else.

The model that ranks your next round is fitted to your own measurements alone. It is built when you ask for a round, from the results you have brought back, and it is thrown away after. Your numbers are never pooled with another organisation's, and no model is shared between organisations. The starting point every lab has before it has measured anything comes from published screens, not from anybody's private results, so nothing you measure improves what another lab is given.

  • Your brief, your materials and your measurements go to the Orcus service and to no outside model. Retrieved papers come from a public index of journal articles.
  • Briefs and results are encrypted before they are stored, with the keys held outside the database.
  • Your organisation sets how long anything is kept, and a scheduled job deletes the briefs, the batches and the feedback that are older than that.
  • Reading papers with a third party model is off unless you turn it on for one run, and what it sends is the text of a published paper and nothing of yours.
  • Every sign in, run and export is written to an audit log your organisation can read.

Your organisation decides how long anything is kept. A record submitted with a feedback note can be deleted on request using the token issued at the time, and an administrator can read the audit log of who did what.

6. Availability and limits of liability

The service is provided during a pilot and may be unavailable, changed or withdrawn. Each organisation has a fixed number of runs. To the extent the law allows, the service is provided as it is, without warranty of fitness for a particular purpose, and is not liable for losses arising from experiments you choose to run. Nothing in these terms limits liability that cannot be limited by law.

7. Changes

Section 3 is rebuilt whenever the measurements behind it change, so it describes the version of the system you are using rather than the version these terms were first written for. We will tell you about material changes to the rest before they apply.

8. Contact

Questions about these terms, about your data, or about anything the service told you: hello@orcus.bio. To have a feedback record deleted, send the token you were given when you submitted it.