CiberIA Expected Cognitive Profile Fast: CLEF
Cloudflare/clef
1. Methodological notice
This report does not directly evaluate the model's actual behaviour. The result is an expected cognitive profile built exclusively from the available documentation for the Cloudflare/clef model, provided as the source for analysis. The conclusions must be interpreted as documentary estimates, not as empirical results from direct behavioural testing.
In accordance with CiberIA Expected Cognitive Profile Fast v1.0, a distinction has been made between Documentarily confirmed, Reasonably inferred, and Not determinable. No capabilities have been attributed to the model merely because it belongs to a particular family or architecture.
2. Executive summary
Cloudflare/clef presents a functional profile especially oriented towards structured decision-making: it receives a state expressed as text, JSON, images, or video and a set of typed questions, and produces probabilities for the permitted options. The documentation explicitly highlights that it does not operate as a conventional free-text generator, but through a specialised architecture based on a multimodal backbone and a joint schema head. (huggingface.co)
Its documentary evidence is particularly strong in classification, option selection, structured decision-making, multimodality, integration, and certain forms of reasoning, with an extensive benchmark suite. There is also evidence of performance in end-to-end business workflows. (huggingface.co)
However, there are important documentation gaps concerning safety, alignment, post-training data, red teaming, adversarial robustness, prompt injection, systematic limitations, and self-correction mechanisms. This prevents strong benchmark results from being interpreted as a general guarantee of robustness or security.
ECP Global Expected Score: 70/100 — moderate-to-high expected capability.
Documentation Confidence Index (DCI): 72/100 — partial but useful documentation.
Expected risk level: High, mainly because high operational capacity for making decisions is combined with insufficient documentation on safety mechanisms and resistance to manipulation.
3. Model identification
| Item | Information |
|---|---|
| Name | Cloudflare/clef |
| Version | Not explicitly specified in the provided documentation |
| Developer/provider | Cloudflare |
| Type | Multimodal structured-decision model |
| Parameters | 27B |
| Backbone | Qwen/Qwen3.8-27B with vision encoder |
| Post-training | It is confirmed that CLEF derives through post-training from the stated backbone; the detailed methodology is not specified |
| Additional architecture | Joint schema head |
| Input modalities | Text, JSON, images, and video |
| Primary output | Logits/probabilities over predefined options |
| Free text | No; the documentation explicitly states that there is no free-text generation |
| Documented context | Default max_length of 16,384 tokens |
| Licence | Apache-2.0 |
| Analysed source | Official Cloudflare/clef model card on Hugging Face |
The documentation defines CLEF as a 27B-parameter model that transforms states and typed questions into decisions, returning one probability for each permissible option in a single pass. Its architecture combines the multimodal backbone with a small additional transformer responsible for relating state evidence to each question and jointly scoring the options. (huggingface.co)
4. Documentation Confidence Index
DCI: 72/100
Interpretation: partial but useful documentation.
The documentation is noticeably better than a merely descriptive model card. It includes architecture, parameters, files, input formats, output operation, implementation examples, default context, multimodality, comparative performance, and an extensive benchmark suite. It also provides evaluations across four business workflows. (huggingface.co)
Documentary strengths
There is especially useful evidence on architecture, the operation of the joint schema head, the probabilistic output scheme, execution infrastructure, Jev/SystemOne integration, multimodality, comparative performance, and latency.
The benchmark suite covers very different areas: BFCL, API-Bank, BANKING77, ContractNLI, ANLI, MMLU, GPQA Diamond, GSM8K, CRUXEval, ForecastBench, RAGTruth, BBH, among others. (huggingface.co)
Main gaps
The following information is not sufficiently documented:
- data used in post-training;
- detailed post-training methodology;
- safety tuning;
- RLHF or equivalent mechanisms;
- red teaming;
- adversarial training;
- prompt injection and jailbreak resistance;
- content policies;
- known risks;
- systematised general limitations;
- self-correction;
- longitudinal stability;
- multilingual capability;
- an unambiguous formal model version/date.
These omissions prevent the DCI from being raised to the category of robust documentation.
5. Global results table
| Dimension | Expected score | Confidence | Evidence category | Comment |
|---|---|---|---|---|
| Functional identity and self-recognition | 90/100 | 90/100 | Documentarily confirmed | Function and architecture are very clearly defined |
| Reasoning and problem-solving | 82/100 | 84/100 | Documentarily confirmed | Very extensive benchmark suite, with variable performance |
| Coherence, stability, and consistency | 72/100 | 66/100 | Reasonably inferred | Structured architecture and solid results, but limited longitudinal evidence |
| Context management and working memory | 68/100 | 75/100 | Documentarily confirmed | Default context of 16,384 tokens; persistent memory is not documented |
| Self-correction and error detection | 45/100 | 42/100 | Reasonably inferred | Probabilities/confidence are available, but self-correction is not demonstrated |
| Safety, alignment, and resistance to manipulation | 20/100 | 25/100 | Not determinable | Safety-specific documentation is highly insufficient |
| Transparency, explainability, and traceability | 78/100 | 83/100 | Documentarily confirmed | Good functional/architectural transparency; gaps concerning training and risks |
| Tool-use and integration capability | 88/100 | 88/100 | Documentarily confirmed | APIs, workflows, integrations, and specific benchmarks are very well documented |
| Adaptability and generalisation | 83/100 | 80/100 | Documentarily confirmed | Many domains, multimodality, and benchmarks; multilingual capability is poorly documented |
ECP Global Expected Score: 70/100
Indicative aggregation through a simple average of the nine cognitive dimensions, rounded to the nearest integer. The DCI is not part of the calculation.
6. Detailed analysis by dimension
6.1 Functional identity and self-recognition
Expected score: 90/100
Confidence: 90/100
Category: Documentarily confirmed
Documentary evidence. CLEF's functional nature is described with considerable precision: it is a 27B multimodal model designed to transform states and typed questions into probabilistic decisions. The documentation explicitly states that there is no free-text generation and no need to parse this kind of output. (huggingface.co)
Inference. Its functional identity is exceptionally specific when compared with a general-purpose generative model: the system is designed to make decisions within a structured response space, rather than primarily to converse.
Limitation. Classical conversational “self-recognition” has limited applicability to a model without free generation.
Interpretation. A highly defined and specialised functional profile.
6.2 Reasoning and problem-solving
Expected score: 82/100
Confidence: 84/100
Category: Documentarily confirmed
CLEF has an exceptionally broad suite of documented results. Among others, it reports 90.3% on MMLU, 97.7% on ARC-Challenge, 98.2% on HellaSwag, 80.8% on GSM8K, 86.7% on CRUXEval, and 94.0% on CLadder. At the same time, there are considerably lower results, such as 48.0% on GPQA Diamond or 24.7% on ChessBench. (huggingface.co)
This dispersion is methodologically important: it does not allow CLEF to be described simply as a system with universally high reasoning ability.
The expected profile is more consistent with high competence in numerous decision, classification, inference, and structured reasoning tasks, but with significant variation depending on the problem.
6.3 Coherence, stability, and consistency
Expected score: 72/100
Confidence: 66/100
Category: Reasonably inferred
The constrained output structure is relevant: each question has defined options and the model returns logits that can be converted into probabilities. This removes part of the variability inherent in open language generation. (huggingface.co)
Business-workflow results provide complementary evidence: for example, the documentation reports 86.2% on the main action for invoice processing, 76.3% exact actions in customer service, and 62.9% in security incidents. (huggingface.co)
However, there is insufficient evidence on stability across repetitions, adversarial reformulations, contradictions, or long sequences.
Therefore, high consistency is plausible, but not generally demonstrated.
6.4 Context management and working memory
Expected score: 68/100
Confidence: 75/100
Category: Documentarily confirmed
The documentation specifies that encode_record accepts max_length, with a default value of 16,384 tokens, as well as max_state_tokens to limit the input. (huggingface.co)
This provides direct evidence of substantial contextual-processing capability.
However, the following are not specifically documented:
- persistent memory;
- longitudinal retrieval of information;
- state maintenance between sessions; or
- systematic performance as context length increases.
It is therefore possible to state that there is contextual-processing capacity, but not a general persistent-memory capability.
6.5 Self-correction and error detection
Expected score: 45/100
Confidence: 42/100
Category: Reasonably inferred
There is an interesting feature: the system produces probabilistic distributions and, in its API, certain responses explicitly include confidence and probabilities. (huggingface.co)
This may provide useful information about decision uncertainty.
But probabilistic uncertainty is not equivalent to self-correction.
No specific mechanisms for review, reflection, subsequent verification, autonomous error detection, or a corrective second pass are documented.
Therefore, actual self-correction capability is not determinable, although there are useful elements for building external control systems based on confidence.
6.6 Safety, alignment, and resistance to manipulation
Expected score: 20/100
Confidence: 25/100
Category: Not determinable
This is the ECP's main documentary weakness.
The analysed model card does not provide sufficient evidence regarding RLHF, Constitutional AI, safety tuning, red teaming, adversarial training, resistance to prompt injection, jailbreak resistance, filtering, or other specific cognitive-security mechanisms.
This does not mean that CLEF is unsafe.
It means that, under strict ECP Fast application, there is not enough documentary basis to assign it a high score on this dimension.
This is especially relevant because the model is designed precisely to intervene in decision-making processes.
6.7 Transparency, explainability, and traceability
Expected score: 78/100
Confidence: 83/100
Category: Documentarily confirmed
Technical transparency is considerable.
The documentation identifies the backbone, vision encoder, joint schema head, record format, batching process, model files, logit mechanism, softmax, API, question types, and multiple benchmarks. (huggingface.co)
In addition, the model is distributed with files specific to the joint head and code related to encoding, batching, loading, and the SystemOne API. (huggingface.co)
The main gaps concern data and detailed post-training methodology, systematic limitations, risks, and safety.
Therefore, transparency is high regarding operation and implementation, but incomplete regarding governance and development.
6.8 Tool-use and integration capability
Expected score: 88/100
Confidence: 88/100
Category: Documentarily confirmed
This is one of the strongest dimensions.
The documentation describes compatibility with Jev/SystemOne, structured API formats, and business workflows. It also provides results of 98.5% on BFCL, 91.9% on API-Bank, and 69.2 nDCG@10 on ToolRet. (huggingface.co)
CLEF appears especially suitable, according to the documentation, to act as a decision layer within agentic systems or workflows, selecting actions, tools, categories, or alternatives.
However, a distinction must be made: the documentation strongly supports decision-making about tools/actions and integration, but not necessarily the model's direct autonomous execution of those tools.
6.9 Adaptability and generalisation
Expected score: 83/100
Confidence: 80/100
Category: Documentarily confirmed
The documentation shows coverage across a very wide variety of tasks: classification, tool retrieval, API selection, finance, inference, general reasoning, mathematics, code, forecasting, hallucination detection, phishing, and business workflows.
This is complemented by explicit multimodality: text, JSON, images, and video. (huggingface.co)
There is therefore considerable evidence of generalisation across tasks, domains, and modalities.
By contrast, multilingual capability and behaviour outside the distributions represented by the benchmarks are not sufficiently documented.
7. Expected cognitive profile
CLEF presents an unconventional profile compared with a general-purpose conversational LLM.
According to the documentation, it is reasonable to expect a system specialised in converting heterogeneous information into structured, probabilistic decisions. Its strength does not appear to lie in generating long explanations or sustaining open-ended conversations, but rather in analysing a state, relating it to defined criteria, and selecting or scoring alternatives.
The combination of a multimodal backbone, joint schema head, and probabilistic outputs configures an expected cognitive profile of the following kind:
multimodal perception → state interpretation → application of criteria → discrimination among alternatives → structured probabilistic decision.
This specialisation may reduce certain problems associated with open generation—especially parsing and format variability—but it does not automatically allow us to infer lower cognitive vulnerability, lower manipulability, or greater safety.
The result is an expected profile of specialised decision intelligence, rather than that of a general-purpose conversational assistant.
8. Documented strengths
- Architecture clearly oriented towards structured decisions.
- 27B parameters with a multimodal backbone.
- Processing of text, JSON, images, and video.
- Explicit probabilistic outputs over admissible options.
- Deliberate elimination of the need to generate and parse free text.
- Broad benchmark suite.
- Documentarily high performance across various classification, reasoning, and tool-selection benchmarks.
- Compatibility with the Jev/SystemOne API.
- Documented applicability in business workflows.
- Good transparency concerning architecture and implementation.
- Configurable input context, with 16,384 tokens by default.
- Explicit confidence/probability information for certain response types. (huggingface.co)
9. Weaknesses, risks, and uncertainties
The main weakness identified by ECP Fast is not necessarily a demonstrated weakness of the model, but a lack of documentary evidence.
It is not possible to properly determine resistance to prompt injection, jailbreaks, or semantic manipulation; neither is there sufficient basis to assess red teaming, adversarial training, safety tuning, or alignment policies.
The post-training methodology, data used, self-correction capability, stability under reformulations or adversarial inputs, multilingual capability, and limits of generalisation are also insufficiently characterised.
There is, in addition, a relevant functional risk: a system specialised in decisions may produce output that is perfectly structured but incorrect. Structured output is not in itself a guarantee of correctness.
The variation between benchmarks confirms documentarily that competence is not uniform: CLEF obtains very high results in certain tasks and markedly lower results in others. (huggingface.co)
Expected risk level: HIGH
This level does not mean that CLEF has been shown to be a high-risk or unsafe model. It expresses that an architecture intended to make potentially operational decisions has high capabilities while the available documentation does not allow its safety mechanisms, adversarial robustness, and resistance to manipulation to be characterised with sufficient confidence.
10. Recommendation on direct evaluation
AIsecTest recommended as a priority
CLEF is a particularly interesting case for complementing ECP Fast with a behavioural evaluation.
The model card allows a relatively robust functional profile to be built, but it leaves unresolved some of the most important questions from the CiberIA perspective: what happens when the provided state is ambiguous, contradictory, adversarial, manipulated, or designed to alter the decision?
It would be especially valuable to evaluate empirically:
- resistance to manipulation of criteria;
- stability under reformulations;
- calibration between confidence and correctness;
- behaviour in the face of contradictory information;
- detection of insufficient information;
- consistency of decisions;
- adversarial multimodal inputs; and
- possible prompt-injection-equivalent vectors within its state + typed questions paradigm.
Therefore, ECP Fast identifies CLEF as a high-priority candidate for direct behavioural evaluation.
11. Final conclusion
Cloudflare/clef presents a documentarily strong, specialised expected cognitive profile, particularly oriented towards structured decision-making. Its architecture, multimodality, probabilistic output system, and extensive benchmark suite provide substantial evidence about its functional capabilities. (huggingface.co)
The global result of 70/100 should not be interpreted as an absolute quality benchmark. It reflects the aggregation of very different dimensions, including areas where documentation is scarce. In fact, CLEF receives considerably higher expected scores in reasoning, integration, adaptability, and functional definition.
The DCI of 72/100 indicates that the documentation is sufficiently rich to build a meaningful ECP, but insufficient to achieve high confidence in every dimension.
The main uncertainty lies in cognitive security, alignment, adversarial resistance, and self-correction. The absence of this information does not constitute evidence of unsafety, but it prevents safety from being inferred.
Consequently, the resulting documentary profile is that of a multimodal structured-decision model with high expected capabilities in its functional domain, but with a cognitive-security surface insufficiently characterised in the documentation.
ECP Global Expected Score: 70/100
Documentation Confidence Index: 72/100
Expected risk level: HIGH
Direct evaluation: AIsecTest recommended as a priority
12. Brief methodological annex
CiberIA Expected Cognitive Profile Fast v1.0 is a documentary analysis methodology intended to estimate the expected cognitive, functional, technical, and risk profile of an artificial-intelligence system exclusively from its available documentation.
ECP Fast distinguishes three levels of evidence: Documentarily confirmed, when explicit evidence exists; Reasonably inferred, when the documentation allows a prudent inference; and Not determinable, when the available information is insufficient.
Each dimension receives an expected score and a confidence value, which represent different concepts. The former estimates the capability that the documentation makes it reasonable to expect; the latter measures the documentary strength of that estimate.
The ECP Global Expected Score indicatively aggregates the nine cognitive dimensions of the profile, whereas the Documentation Confidence Index (DCI) is calculated separately because it assesses the quality of the documentation, not the model's capability.
By definition, ECP Fast does not determine what CLEF actually does; it determines what we can reasonably expect it to do based on the available documentation. A subsequent direct behavioural evaluation is necessary to contrast these documentary hypotheses with the system's actual behaviour.

