Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Positive Sentiment Can Still Contain Criticism: Testing Jev Alongside ABSA
Latest   Machine Learning

Positive Sentiment Can Still Contain Criticism: Testing Jev Alongside ABSA

Last Updated on September 22, 2026 by Editorial Team

Author(s): Massimiliano Geraci

Originally published on Towards AI.

Positive Sentiment Can Still Contain Criticism: Testing Jev Alongside ABSA

A recovered local baseline, complementary judgments, and the limits of an early shadow replay.

Positive Sentiment Can Still Contain Criticism: Testing Jev Alongside ABSA
Figure 1. A positive label can coexist with separate praise and criticism judgments. This is a conceptual illustration of the diagnostic, not a human-validated interpretation.

The local classifier called the passage positive. Jev returned both praise and criticism.

In one anonymized case from our ORCA Signal experiments, a selected Jev draw placed praise at 2.32 and criticism at 0.48, on separate zero-to-three rubrics. The context was not truncated, and the application checks admitted that output. Our existing ABSA client returned positive for the same passage.

That is an interesting result. It is also easy to oversell.

There was no independent human label establishing whether Jev had correctly identified the reservation. We had an observed pair of outputs, not a verdict about which model understood the passage better. My interest is in the application question: could those additional judgments help someone interpret an answer that a single polarity label summarizes as positive?

What each component was measuring

This follows the broader Jev pilot I described in Jev Beyond the Demos: Testing System One Models in ORCA Signal. There, we investigated bounded decisions around claims, evidence, sentiment, and possible writing-workflow gates. This follow-up narrows the question to a sentiment sidecar.

Our local comparator used this DeBERTa model for aspect-based sentiment analysis (ABSA). In this replay, its inputs were text plus brand, and the client produced brand-global polarity using its existing label rule. That rule is not necessarily the class with the highest probability. We preserved it.

ABSA can address aspect-specific sentiment. We did not configure this comparator to annotate the particular aspect definitions supplied to Jev. That distinction belongs to our experiment, rather than to a claim about ABSA’s capabilities.

ORCA Signal separates category-level sentiment from mention counts and preserves missing category assignments. This A. O. Smith product dataset illustrates the workflow; it is not the Jev–ABSA replay reported here.

Jev’s public interface lets us supply state and ask independently defined, typed questions. We used that interface to separate praise, criticism, attribution, and scope. Its aspect-specific arm also received an aspect definition. The Score primitive evaluates an ordered rubric; our zero-to-three outputs are rubric-based intensities, not ABSA class probabilities.

Sharing text does not make those questions or scales equivalent.

Figure 2. Sharing a passage does not align the target, question, or scale. The aspect-specific arm also supplied an aspect definition.

Recovering the missing comparator

The previous replay had already produced Jev receipts, but our local ABSA attempt had failed before supplying a usable baseline. A native dependency could not load. We recovered the comparator in an isolated CPU process, correcting the temporary dependency mount while keeping the original client, model, tokenizer, and input plan. All six model and tokenizer file hashes matched the earlier attempt.

The recovery produced 34 valid first evaluations and 34 identical warm cache hits. Those cache hits are repeat invocations, not another 34 independent observations. We reused the 216 existing Jev receipts, with no new Jev calls. Recovery added no provider charges; local CPU time still has a cost.

This established a local shadow comparator. It did not exercise the remote runtime, database integration, full preprocessing and routing pipeline, calibration, or human adjudication.

Count the sources before the requests

The real material consisted of 17 original answers from an evidence-selected legal-services market corpus. Each answer supplied two context selections.

The canonical selector retained a contiguous original-text envelope around selected evidence, including intervening text. The legacy selector used an existing brand-local window, with its sentence and spacing transformations. Neither is automatically the correct context for every judgment.

Download the Medium App

Seventeen answers therefore became 34 contexts. Jev evaluated two scopes, global and aspect-specific, with three draws per context: 204 real-data requests. Four synthetic controls, repeated three times, supplied the other twelve receipts. The resulting 216 requests are not 216 independent real examples.

The ABSA tokenizer exposed another boundary:

Seven text-and-target pairs exceeded the local limit. We retained their receipts but excluded them from the matched matrix: truncation could change what the comparator actually evaluated.

The 27 remaining contexts generated 81 global Jev draw records. Repetition still does not create new source answers. Also, the twelve canonical and fifteen legacy contexts are different populations, so comparing their percentages would not isolate a selector effect.

A positive sentiment label can accompany qualified language: this persisted response describes a redundancy benefit while noting the absence of a built-in controller. The excerpt illustrates why context inspection matters; it does not establish a classification error or verify the technical claim.

Only ten source answers had both selections untruncated. Within that paired population, two ABSA labels changed: one neutral to positive, the other positive to neutral. Without human labels, neither direction establishes an improvement.

What a sidecar might expose

The opening case makes the design concrete: retain the existing positive label and separately inspect Jev’s praise, criticism, attribution, and scope judgments. We have not validated a new aggregate score, and there is no reason to invent one merely because multiple numbers are available.

In ORCA Signal’s Brand Performance Market analysis, that could help a reviewer distinguish an adopted reservation from a criticism the answer merely reports or rejects. In objection analysis, the next question is whether a reservation changes selection. A polite statement about a missing mandatory capability can matter more than emphatic dissatisfaction.

Those are separate hypotheses. The recovered ABSA replay did not establish objection severity or decision consequences. It also does not authorize changes to the Truth Analysis Index, our evidence-oriented framework for examining whether claims are supported by their references. Complementary sentiment judgments need their own validation.

The application checks remain consequential. Jev’s independent fields can compose into an incoherent result. Our earlier canonical replay still contained ten incoherent aspect-specific draws, compared with three under the legacy selection. Recovering a brand-global comparator cannot resolve those attribution and scope problems.

An insufficient context stays insufficient. A reported opinion does not become the assistant’s adopted opinion. An incoherent combination stays under review. None of those states should disappear into a neutral label or zero intensity.

Figure 3. The candidate sidecar preserves the existing polarity and adds separate judgments. Application checks can withhold them; no new aggregate sentiment score has been validated.

The evaluation we need next

The corpus contains no negative ABSA labels and no real-data Jev negative-only orientations. It is evidence-selected and small. It cannot tell us whether this sidecar handles strong criticism across broader markets.

I would build the next gold set around independently annotated passages, with brand and aspect targets stated explicitly. Annotators would label net polarity, praise and criticism, speaker attribution, and scope separately, before seeing model outputs. The gold set is a planned evaluation asset; the current review packets remain unannotated.

We also need stronger negative cases, reported and rejected criticism, ambiguous targets, and incomplete passages. Related answers and their context variants should stay in one data split. Rubric development should use one partition; the final evaluation should use held-out source families.

Then compare reviewer performance with the original polarity alone and with the proposed sidecar. Do the extra judgments improve correct interpretation? Do they expose useful reservations without adding misleading ones? What happens to total review time?

Cost and latency belong to that complete path, including context selection, inference, application checks, and escalation. A warm local cache hit compared with a fresh HTTP request would answer the wrong performance question.

For now, we have recovered a baseline and identified a plausible addition to it. The next experiment is to establish whether a reviewer can use that addition to make a better, evidence-supported judgment.

Figure 4. Independent labels, held-out source families, and a reviewer utility study are the next evaluation. None has yet established a benefit for this sidecar.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.