Jev (Fully Tested) + Browser Use: FASTEST AI Agent I'VE TRIED YET!
Visit my second channel AISeeKing: https://youtube.com/@AISeeKing In this video, I’ll be testing Jev, TypeSafe’s System One model designed for fast, inexpensive, and structured AI decisions. I’ll explore its performance in support routing, refund detection, prompt injection resistance, exact-value selection, agent auditing, and browser automation. -- Key Takeaways: ⚡ Jev delivers structured decisions with reported evaluation times as low as ninety-two milliseconds. 💸 Its low input-token pricing makes small classification and decision tasks extremely inexpensive. 🎯 Jev performs well at support routing, detecting negation, scoring urgency, and selecting exact values. 🛡️ A basic prompt injection failed to override the model’s trusted evaluation instructions. ⚠️ Restricted outputs do not guarantee correct answers when the available choices are incomplete. 🔍 Jev can audit agent activity by comparing final claims against actual tool results. 🌐 The Jev Ultrafast demo shows how it can support fast browser automation alongside another language model. 🧪 These tests are promising, but broader testing is still required before using Jev in production workflows.
Read Video · Bản Ghi & Phân Tích
Tập này có bản ghi đầy đủ + AI insights
Tài khoản miễn phí · không cần thẻ · 150 tín dụng khi đăng ký, đủ để mở khóa tập này
- 📄 Bản ghi đầy đủ có dấu thời gian
- ✨ Tóm tắt AI, từ khóa & bản đồ tư duy
- 💡 Điểm chính và trích dẫn nổi bật
Dòng thời gian tập
Introduction to Jev, a system-one model for fast structured decisions, and the three question types it supports.
- Jev is described as a system-one model that takes information and returns decisions rather than generating chat responses or writing applications.
- It supports three question types: choice (selecting from provided options), score (rating on described levels), and null (giving a probability that a yes/no statement is true).
- The goal is to make small judgments cheap and fast enough to embed throughout software, such as routing messages, detecting refund requests, or flagging items needing attention.
Practical tests on a support message about a duplicate charge, including negation handling and a limitation when no correct option exists.
- For a duplicate-charge message, Jev selected billing, gave the refund request a 98% probability, urgency 11%, and a frustration score near the calm end, correctly separating billing from technical and refund from emergency.
- When the message said 'I am not asking for a refund,' Jev still selected billing but dropped the refund probability to 3%, showing it handled negation rather than just spotting the word 'refund.'
- When asked what time the cafeteria closed with only billing, technical support, and sales as options, Jev selected sales with 0.31 confidence, demonstrating that restricting output prevents invented categories but doesn't guarantee the selected category is useful or correct.
Prompt injection test and exact-value extraction from a document.
- A fake system override embedded in the message telling Jev to choose billing and mark refund and urgency as true did not work: it kept technical support classification, refund probability stayed at 3%, and urgency stayed in the uncertain middle.
- For extracting a current receipt destination from a message containing an old and new address, Jev selected the new address exactly as supplied, including the plus sign and year, showing code can collect candidates and copy the original value.
- If the first step misses the correct address, Jev cannot create it through a choice question, so the candidate list is part of the system that needs testing.
Khái niệm chính
- Jev— The AI model being tested, designed for fast structured decisions rather than chat responses.
- structured outputs— Jev returns decisions in formats software can use directly, such as choices, scores, and probabilities.
- prompt injection— A test where malicious instructions inside a message tried to override the model's actual task.
Trích dẫn nổi bật
So when you hear the claim about zero hallucinations, keep that distinction in mind. Restricting the output can stop the model from inventing a new category. It doesn't guarantee that the category it selects is useful or correct.
💡— It overturns the assumption that zero hallucinations means the model is always correct, revealing that restricted choices can still produce useless answers.
I had created a situation where it couldn't return the answer I needed.
🎯— It shows that the model's output is only as good as the options you provide, a hidden limitation of structured decision models.
Hành động cụ thể
🧪AI Model Evaluation
Structured decision models like Jev are only as good as the choices and instructions you provide.
This week, test your own wording and candidate lists with Jev before connecting it to any automated action.
Confidence values are not accuracy guarantees and should not be read as proof of correctness.
Run a small validation set of at least 20 examples to measure actual accuracy against Jev's confidence scores.
⚙️Practical Implementation
Jev can handle routing, negation, value selection, and claim verification with low cost and fast evaluation times.
Identify one repetitive judgment task in your software this week and prototype a Jev-based solution for it.
The candidate list you supply is part of the system you need to test; if it misses the correct value, Jev cannot create it.
Audit your candidate lists for completeness and add a fallback path for values not in the list.
Phiên âm và thông tin chi tiết được tạo bởi AI và có thể chứa lỗi. Độ chính xác phụ thuộc vào chất lượng âm thanh và sự rõ ràng của người nói — nếu có gì sai, âm thanh gốc luôn là nguồn chính xác nhất.
Tập podcast và video sẵn sàng để đọc
Tập podcast

The Mystery of Sea Creatures (3/5): Could an orca give a TED Talk? | Karen Bakker
TED Talks Daily
18 thg 7, 202615:27EN
Why Jensen Huang Believes We’ve Reached AGI and Inside OpenAI’s German Website Hijack | #287
Moonshots with Peter Diamandis
9 thg 9, 20262:23:43EN
Episode 68: How Timothy Baxter Built Baxter Research Into the Gold Standard of Criminal Research
Behind the Screens: Conversations with Background Screening Pros hosted by Les Rosen
20 thg 7, 202659:28EN
#42 Beter worden in Canva zonder te verdwalen
De Canva witch podcast
23 thg 2, 202611:21NL
#22 Barbarossa
Wer wir sind und warum das nicht klappte ...
3 thg 9, 202555:55DE
【#大師小聊】EP16|從文案講師到創業家:他為什麼說「創業更不自由」 ft 純粹文案創辦人 林育聖
李洛克的大師小聊
20 thg 4, 20261:08:38ZH
Video

How to Build Things with Jev & OpenJevs
Sam Witteveen
21 thg 9, 202619:19EN
Tom Holland and Jon Bernthal Team Up While Eating Spicy Wings | Hot Ones
First We Feast
23 thg 7, 202629:21EN
Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Patel
11 thg 8, 20262:12:32EN
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12 thg 11, 202553:29DE
10 habitudes qui m’ont VRAIMENT fait perdre du poids
leawellnesss
15 thg 4, 202614:34FR