Intro to Content Defined Chunking
Building a basic search experience with Postgres
Building search functionality in products is a common task. Many solutions exist to solve this problem already. OpenSource tools like opensearch and mellisearch are some examples that are very commonly used. Using a 3rd party tool to build a “full-text search” is a good bet if you have a lot of data (i.e. a lot of users).
The goal of this post is to have a look at some in-built tools to build a minimalist search feature when your data is backed by a Postgres (or any SQL) database.
The workload-first approach to CPU innovation
Sponsored Post: As in so many other aspects of life, not all compute workloads are created equal – they need a more subtle approach to getting the best
Why Audit Logs Are Important
Administrators and compliance teams can use audit logs to investigate user actions, spot suspect activity and adhere to regulatory frameworks.
Zero-copy I/O for ublk, three different ways
The ublk subsystem enables the creation of
user-space block drivers that communicate with the kernel using io_uring. Drivers implemented this way show
some promise with regard to performance, but there is a bottleneck in the
way: copying data between the kernel and the user-space driver's address
space. It is thus not surprising that there is interest in implementing
zero-copy I/O for ublk. The mailing lists have recently seen three
different proposals for how this could be done.
Rules as code for more responsive governance
Using rules
as code to help bridge the gaps between policy creation, its
implementation, and its, often unintended, effects on people was the
subject of a talk by
Pia Andrews on the first day of the inaugural
Everything Open conference in Melbourne, Australia. She
has long been exploring the space of open government,
and her talk was a report on what
she and others have been working on
over the last seven years. Everything Open is the successor
to the
long-running, well-regarded linux.conf.au (LCA); Andrews (then Pia Waugh) gave the opening keynote at LCA 2017 in
Hobart, Tasmania, and helped organize the 2007 event in Sydney.
Four things you should know about Rules as Code
7 Hardware Devices for Edge Computing Projects in 2023
Edge computing can ensure things keep functioning when reliability is vital and network connectivity can’t be guaranteed.
Visible APIs get reused, not reinvented
How open API specifications can help developers—and computers—understand your APIs.
FLEDGE - Chrome Developers
A proposal for on-device ad auctions to choose relevant ads from websites a user has previously visited, designed so it cannot be used by third parties to track user browsing behavior across sites.
Partnering with Fastly—Oblivious HTTP relay for FLEDGE's 𝑘-anonymity server - Chrome Developers
We are improving Chrome’s privacy measures by partnering with Fastly to implement the 𝑘-anonymity server for FLEDGE. With data being relayed through an OHTTP relay in this implementation, Google servers do not receive the IP addresses of end users. The 𝑘-anonymity server is an incremental step towards the full implementation of FLEDGE.
You can now run a GPT-3 level AI model on your laptop, phone, and Raspberry Pi
Thanks to Meta LLaMA, AI text models have their "Stable Diffusion moment."
VEX: Standardization for a Vulnerability Exploit Data Exchange Format
VEX documents, a companion to software bills of materials (SBOMs), describe how and whether a piece of software is affected by a certain vulnerability.
Building a Media Understanding Platform for ML Innovations
The media understanding platform serves as an abstraction layer between Machine Learning (ML) algos and various applications.
DOE Wants A Hub And Spoke System Of HPC Systems
We talk about scale a lot here at The Next Platform, but there are many different aspects to this beyond lashing a bunch of nodes together and counting
Time is Code, my Friend
For several years, the fast startup time of an application or service in Java has been a highly discussed topic. According to the authors of numerous articles, it seems if you don’t apply som…
Naming conventions in programming – a review of scientific literature — Makimo – Consultancy & Software Development Services
Every programmer faces the problem of finding good names to communicate the author’s intent. See how science answers this question and get actionable insights.
Revisiting stamps for email
I started agitating for this in 1997 and wrote about it in 2006. The problem with the magical medium of email is that it’s an open API. Anyone with a computer can plug into it, without anyone…
File Systems in Operating System - GeeksforGeeks
A Computer Science portal for geeks. It contains well written, well thought and well explained computer science and programming articles, quizzes and practice/competitive programming/company interview Questions.
Embedded and cloud Elixir for grid-management at Sparkmeter
A case study of how Elixir is being used at SparkMeter.
False Sharing versus Perfect Placement - Marc's Blog
Parsing JSON Really Quickly: Lessons Learned
Daniel Lemire talks about the lessons learned while writing the fast JSON parser, simdjson. One of the most important lessons is the importance of a nearly obsessive focus on performance metrics - to constantly measure the impact of the choices.
Irregular expressions - tavianator.com
Husky: Exactly Once Ingestion and Multi Tenancy at Scale
A closer look at storage routing in Husky, Datadog's third-generation event storage system.
Signal K » Welcome
Signal K is about publishing a common modern and open data format for marine use. A format for the modern boat, compatible with NMEA, friendly to WiFi, cellphones, tablets, and the Internet. A format available to everyone, where anyone can contribute.
How Decoupling Can Help You Write Better Software
Decoupling components enables you to deal with performance issues, cost optimization and feature development separately for each service.
What Is ChatGPT Doing … and Why Does It Work?
Stephen Wolfram explores the broader picture of what's going on inside ChatGPT and why it produces meaningful text. Discusses models, training neural nets, embeddings, tokens, transformers, language syntax.
You Probably Shouldn't Mock the Database – dominikbraun.io
To keep unit tests fast and isolated, the data access layer is often tested using a mock of the database. But are unit tests and mocks actually a good choice?
Programming Language over Data language
The personal website of JT Archie. Includes a blog, work ethic, and projects they have worked on.
PEAK:AIO provides HPC-level performance for AI with NAS simplicity and low cost – Blocks and Files
UK startup PEAK:AIO has rewrittten some of the NFS stack and LInux RAID code to get a small 1RU server with a PCIe 5 bus sending 80GB/sec of data to a single GPU client server for AI processing. Three cheap servers doing this would send 240GB/sec to the GPU server, faster than high-end storage arrays […]