Fara 1.5-9B W8 baseline evaluation

This repository preserves the exact Fara 1.5-9B checkpoint used for both W8 DOM browser-agent baseline evaluations. The checkpoint was not trained or fine-tuned for either evaluation. Its files are an exact mirror of the pinned Microsoft checkpoint.

Common checkpoint and serving setup

Component Value
Base checkpoint microsoft/Fara1.5-9B
Base checkpoint revision 1a93677cd89d5601bc2ed759791e981f3a520032
Fara source revision a675d6d61c41c47ae87bacefeab22caad18e3e84
Model dtype BF16
KV-cache dtype BF16 through vLLM auto
Maximum model length 32,768 tokens
Tensor parallel size 1
Runtime vLLM 0.19.1, PyTorch 2.10.0, Transformers 5.6.2
Hardware One NVIDIA H100 80 GB HBM3
GCP instance w8-fara-9b-8115f370-head-3ibtw2ah-compute
Deployed DOM adapter SHA256 6184fba5dee3cae079b20fe1abb1812e589e68b97737a28dac94b16e7f5c5811
Deployed serving script SHA256 1d1c9c60280b15a73d633717776aafd23d087f6af447524ef9cc8315e6d70fb3

The DOM serving integration translated the recorder's OpenAI Responses contract to the Fara inference endpoint. It did not modify the model weights.

Evaluation 1: 90-task baseline

Measure Result
Selected tasks 90
Passed 6
Verified failures 60
Execution errors 24
Infrastructure errors 0
Pass@1 6.67%
Pass rate among verified tasks 9.09%

Evaluation 1 provenance

Component Exact source
Evaluation artifacts WootzappLab/fara-baseline-artifact
Immutable artifact commit 30840bf827a58cfd0bdc8c86fa730ab32960448c
Task dataset WootzappLab/cua-bench
Dataset revision 745eebb8f07abb92fdaae356bb6bee3a1f2975af
Recorder, browser harness, environments, and verifier tokenbender/w8-cua-bench
Harness revision 66037db6bae4b00a2b5860aa08ebb69b559e6871
Evaluation date 2026-09-25

Execution errors remain in the denominator. The artifact repository contains the raw recordings, logs, rubrics, verifier outputs, structured tables, and integrity manifests.

Evaluation 2: 40-task baseline

Measure Result
Selected tasks 40
Passed 8
Verified failures 24
Execution errors 8
Infrastructure errors 0
Pass@1 20.0%
Pass rate among verified tasks 25.0%

Evaluation 2 provenance

Component Exact source
Evaluation artifacts WootzappLab/Fara-40-baseline-artifact
Immutable artifact commit 0a0a615d2d04ad072972f657cf6a5c27cbd54b49
Results bundle SHA256 43f4ed88b79d9c0abe7a733ac6499b0000364e3f1f676a29320e83d1682e081b
Exact task snapshot Included in the artifact repository
Task snapshot SHA256 745f113b9c45f1a8dc3581ebc671435a3eafdfdbfb38bcaa6df1097ec10ab2e3
Dataset repository revision Pending separate publication
Recorder, browser harness, environments, and verifier tokenbender/w8-cua-bench
Harness revision 09976c86454f9a87e2fa6a3a3589d2e4d24f180e
Infrastructure revision eb5dc3d9e0cbfbe389334b0600002f8db3a795dd
Evaluation date 2026-09-28

Execution errors remain in the denominator. The exact 40-task input snapshot is bundled with the artifacts while its separate dataset publication is pending.

Repository contents

  • Model configuration, tokenizer, processor, chat template, and Safetensors weights copied from the pinned upstream revision
  • CHECKPOINT_FILES.json: upstream Git blob identities and LFS SHA256 hashes for the checkpoint files
  • MANIFEST.json: machine-readable checkpoint provenance
  • PROVENANCE.md: checkpoint revision and experiment-boundary receipt
  • CHECKSUMS.sha256: hashes for the repository's reproducibility metadata

Load the checkpoint

import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "WootzappLab/fara-9b-baseline-eval"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

License and attribution

The mirrored base checkpoint is released by Microsoft under the MIT License. See the upstream Fara 1.5-9B model card for its original documentation, intended use, and limitations.

Downloads last month
23
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WootzappLab/fara-9b-baseline-eval

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(1)
this model