Monday, August 5, 2019

Quality of Time

Atomic Rules has been involved with the "Quality of Time" since we began back in 2008. Quality of Time to us means a specified precision and accuracy distributed coherently across an entire domain. If you are only counting milliseconds, this isn't such a chore. But as we move orders of magnitude to microsecond, or even nanosecond resolution; there numerous pragmatic challenges. 

Atomic Rules TimeServo product, introduced in 2016, first commercialized the idea of an FPGA system timer. The ability to coherently deliver time to any component within an FPGA, in the clock domain that time is desired in, is a common need which TimeServo fills.

We purposefully divorced the business of keeping time from the process of protocol-based time transference such as 1588v2/PTP and White Rabbit. TimeServo keeps time... period. There are both closed- and open- loop modes for the phase-accumulator we term BigPhase to operate in. For the closed-loop Digital Phase Locked Loop (DPLL) mode, we turned to mathematician Steve Gabriel to get the inherent structure and corresponding coefficients just as our customer's desired. Most TimeServo users stick with Steve's tradeoffs for settling-time and stability. But we left them programmable from the control-plane just in case.

It's really important to understand your requirements when it comes to Quality of Time. And as you press down below microsecond accuracy towards nanosecond accuracy, things can get difficult quickly. There's a lot of good spec work in the time-centric IEEE standards, and in how organizations such as the UNH IOL can verify and validate performance. We believe there are way too many devices on the market touting "sub-nanosecond accuracy"  - when that's far from the case.

Wednesday, January 10, 2018

The Role of an Advisor

For the past nine years, Atomic Rules has been evolving. Evolving from a single engineer with a singular vision, to an amalgam of a dozen engineers with a team vision. Along the way we've had a lot of help. Help from technology firms like ARM and Xilinx producing the canvas that we draw on. Help from businesses like BittWare who represent our IP products with a worldwide reach. And of course Bluespec Inc, for Bluespec SystemVerilog (BSV), so that our codes would be correct and concise, in a way other Domain Specific Languages can only someday hope to match.

I started Atomic Rules back in 2008 solely offering FPGA design services. We picked up some great clients along the way. A few years in, the idea of producing IP products took hold. Well chosen, customer-driven IP cores would be the future of the company. Not for the squeamish, we knew going in that production IP would require an investment in verification and support significantly above that of just FPGA design services. We got that right.

As of January 2018, Atomic Rules now has three full-production IP cores on the market. Available through Xilinx, available through BittWare, and available directly from Atomic Rules. Hooray for us! With such deep and solid investment in products, tens of engineer-years in the making; it’s time for Atomic Rules to pivot and focus even more on sales, customers and product support; and less intensely on speculative R&D.

We will, of course, keep on our existing services clients, continue our relationships with the dozen engineers who make up Atomic Rules, and continue to support and develop new products and technologies. As such, we have a new Chief Operating Officer to do much of the day-to-day triage; and will soon announce a new Engineering Manager. It's all good!

This week I started a full-time position as VP of Research and Development at an exciting startup, Skreens Entertainment, and thus have relinquished my position as CTO of Atomic Rules. Skreens understands that Atomic Rules shall remain a going and growing concern, and I will stay on in an advisory role. And while my daily (and nightly) passion is now making Skreens all that it can be; I'm fortunate that I get to watch the Atomic Rules rocket climb up to orbit as well.

Saturday, June 3, 2017

FPGA with AXI plus GPP with DPDK equals Arkville

And just like that, the passion that has taken so much effort and focus for so long, is launched. The launch this week of Arkville was, for Atomic Rules, a very big deal. By far our most ambitious product to date. We had a lot of help, and really could not have done it alone. The support of The Linux Foundation and the DPDK Community was vital. They ran a must-read terrific interview that really sums it up nicely.

Are we done? No way! Work is already underway for the 17.08 release. Check back here in August.

Tuesday, May 2, 2017

Hoplite Comes of Age

Jan Gray's groundbreaking work on Hoplite has spread far and wide. Besides Jan, nobody has grabbed that baton and run farther with it than Nachiket Kapre. At FCCM2017 earlier this week, we were treated to not one, but two, presentations extending Hoplite in exciting and new directions.

FPGA overlays are a peculiar thing. The proliferation of spatial heterogeneity in reconfigurable computing devices (Trimberger FPL2007, et al) has effectively killed off the promise of a relocatable circuit module.  Overlays fight back, in part, by delivering an abstraction that helps hide the spatial speed bumps.

Kudos to Jan, Nachiket and others who have advanced this work. They are delivering value to the community that may not be fully appreciated until we have our second 25 year retrospective.




Sunday, March 19, 2017

The Road to Arkville

For over a year now, five engineers have applied their passion on the journey of birthing a new product. Arkville is a big deal for Atomic Rules. Coming 18 months after our first product, the UDP Offload Engine, Arkville will be a whole new bag. We're betting that there is significant value at the intersection of two key technologies: DPDK and AXI.

We dont expect this to be an easy journey. Why? Let's take DPDK and AXI one by one.

DPDK is the de facto Linux Kernel Bypass mechanism evolving now for over a decade for use when you need to do useful work on multiple cores without the kernel stack getting in the way. Most DPDK users work with merchant ASIC-based NICs or virtualized NICs. Despite the "Data Plane" in DPDK, not all users see it as just an I/O mechanism. There's more to it.

Also over the past decade, AXI, a proper sub-set of AMBA, has emerged in the FPGA world as a standardized hardware interface and API. If you are building a Green-box fleet of RTL accelerators to go inside your FPGA; you probably want to use AXI for the plumbing. If not, you are probably writing gaskets with a throughput and latency tax.

Atomic Rules understands that both the DPDK GPP software world and the AXI RTL gateware world are two different things. And experts in one are seldom experts in the other. The guiding light for Arkville design has been to take the established DPDK APIs and ABIs and implement a DPDK- and AXI- aware packet conduit between GPP and FPGA. Conceptually abstract, but physically (at product launch in May) to be PCIe. We're excited about all of this, and if you are too, please get in touch. Since you made it this far, here's a hidden link on our site about Arkville in alpha-testing.


Monday, January 16, 2017

Paced Packet Player

We employed our Esopus Creek technology to build a DPDK Paced Packet Player (PPP), shown here inside a Dell R730 server. Four independent 100 GbE ports, line rate, ns accurate, no waiting. Red arrows show the added 12V supply used in this instance. Watch for more Esopus Creek in our upcoming Arkville product offering later this year.

Wednesday, December 2, 2015

Exciting Times

It’s almost 2016. 20nm FPGAs are here and 16 nm are coming. You can buy an inexpensive 128-port OpenFlow switch/router with 10/25/40/50/100 GbE ports. Intel is roaring back into networking with Red Rock Canyon (RRC). Opportunities abound at the endpoints of these commodity communications networks!

An exciting time for Reconfigurable Computing (RC):
  • 20 nm FPGAs are in production
  • 20 nm FPGAs have 25 Gbps+ capable SERDES
  • 25/50/100 GbE technology-enablers
  • Able to serve 10/40 GbE by running SERDES at lower rates
An exciting time for Software Defined Networking (SDN):
  • The P4 language is rapidly evolving
  • Game-Changing Broadcom Tomahawk
  • Game-Changing Intel Red Rock Canyon (RRC)
An exciting time for Atomic Rules:
  • Timing closure at 400 MHz through architecture
  • A UDP/UOE product, first mover for 25 GbE
  • 10/25/100 GbE L2/L3 stacks using production IP

Saturday, April 19, 2014

Dealt a new Deck

We are so excited with the release this week by Xilinx of Vivado 2014.1 . Having been involved since the dark-ages with the underpinnings of this world-class CAD tool for FPGA, we're always eager to see what's new and improved. With almost a dozen 28 nm designs under our belt, we made quick work of porting designs working on platforms like the KC705, VC709, and ZC706 from 2013.4 to 2014.1. We love that we can script in Tcl, or click in the GUI. We like to see both work equally well. We need more time to have hard quantitative evidence, but across the board we are seeing 10~20% placer run time reductions, with no loss of Quality of Results. Not sure where this gain is coming from, but we will measure and find out. This is crazy good. Thanks Xilinx!

Altera is looking to answer with their salvo of Quartus II 14.0 soon, so stay tuned. Altera does have some interesting silicon announced down the pike; and we're happy to bring our RTL to whatever devices our clients are using. Still, from a do-it-all CAD tool perspective, we lament that Quartus lacks a built-in functional simulator like Vivado. Oh well... separation of concerns!

What a world it would be if Xilinx and Altera got out of the CAD tool business and just focused on world-class FPGA silicon, leaving the EDA tools to the EDA tool vendors. Or to the open source community! It's nice to dream. Still, with tool suites like Vivado that build in aspects like IP Integrator, and IP Packager, we're glad (for now) that the FPGA vendors are in the EDA game.

Sunday, March 30, 2014

Functional Correctness Quickly

A saying we use around here that has evolved from axiom to dogma is this:

“We shall achieve Functional Correctness Quickly and Performance Correctness Iteratively”

In a fast-moving world, this has substantial value to our clients who are often looking for us to operationalize a new platform or algorithm. They turn to us to provide a “show-me” demo under tight time and cost constraints. That near-in demonstration is rarely the end goal. Frequently they want to tune for lower latency, more throughput, better energy efficiency, lower gate area, and add features needed to address their continuously evolving market requirements. We live for this. This is our passion.

How do we do this? In our world, where reconfigurable computing intersects with the business of complex concurrency, we have some things dictated to us and we have places where we have some room to choose.

In the dictated-to-us department we have the FPGA back-end tools for synthesis and implementation. Xilinx dictates Vivado, Altera dictates Quartus. Tabula dictates Stylus, and so on. The vendor tools all have their strengths and weaknesses; and they all provide the capability to go from RTL to bitstream. We have our preferences here, but our clients come first. If they choose an Efinix device, we will use the flow thus dictated to us. Fine.

It is in the higher-level choice of Domain Specific Language (DSL) where our choice counts the most, especially as it relates to “Functional Correctness Quickly”. In our agile sprints to quickly achieve first-functional code, we have used Bluespec SystemVerilog (BSV) for nearly a decade.

The fire-drill to first-functional often goes something like this: We rough out a relationship of various modules and their hierarchy. Each module provides an interface. Without getting caught up in the implementation that goes in these modules, we are free to think solely of the interaction patterns between them. We go through this healthy IP birthing-phase where we may anthropomorphize the module on behalf of the designer. “Shep’s TagServer is in the business of producing an ordinal sequence of tags to a plurality of clients”. In this first step we are very much concerned with what each module does, not how it does it, and what sort of interaction patterns are needed with other modules. We may even assign abstract types to the interfaces because we don’t know yet what is needed.

Bluespec System Verilog (BSV) is exceptionally well-suited for this kind of scrum. We work in either the command-line or GUI-based Bluespec Development Workstation (BDW) mode rapidly crafting Types, Interfaces and Modules. Feedback from design entry to viewing compiled results comes in just seconds; an edit compile debug cycle that is no laggard to the best software development techniques. When we’re done with this step we have the rough-structure of our design in place, which I like to call scaffolding.

Getting to first-functional, the Functional Correctness Quickly part, requires us to provide an implementation for our modules. Because we spent the first step compartmentalizing what does what and how modules communicate, this second-step goes more-quickly as we already know precisely “what business” every module is in. This focus not only let’s us code more concisely, but is also well-suited for quick, on-the-fly, DUT testbenches that add a built in verification aspect to the design process. We will usually code both standalone Bluesim simulations, which are C-based bit- and cycle- true simulations that can be seen in less than a minute; as well as Verilog-RTL simulation which we can run on a Verilog simulator (rarely) or push through to the FPGA and run on hardware (often). Because all BSV code synthesizes there is no special step or flow for this. The benefit is that we see feedback from the FPGA vendor tools for area and Fmax early in the module component development process.

With our collection of modules implemented, we push through the entire design to FPGA. Since the individual modules have been tested, and the interconnection of modules are tested and standardized, it is infrequent that the system-level composition exhibits unexpected behavior. In the rare case where a behavior is not understood, we bisect until we locate the issue. But most of the time, ta dah, we have our first functional code running. Functional Correctness Quickly!

But we may not be done. We will measure the metrics and check against the current requirements to see if we are. It is common to see requirements change. A design that does not tolerate change well is said to be brittle. BSV designs are the opposite of brittle. Let’s say the TagServer we described earlier needed to be embellished to produce a tag type that was not just a unsigned integer, but more-elaborate data structure. Fine! Have at it! With a few tactical changes the modules that produce and consume tags can be adapted. Changes that would be gut wrenching in VHDL or Verilog are almost effortless. And conceptually you are empowered to change the “what” (Types) and how (Rules) separately. Bottom line here is that you get your Performance Correctness Iteratively!

We acknowledge that there may be other ways to realize our agile practice of racing to first-functional and then continuously improve. However we know of no better EDA tool than BSV, when dealing with parallel communicating sequential processes, what we call Complex Concurrency, so common in FPGA.

Sunday, February 23, 2014

The Ultimate Digital Designer’s Assistant

For most of the past decade we’ve had the choice of implementing the solutions to our client’s needs in RTLs, such as VHDL, Verilog, or SystemVerilog; or in Bluespec SystemVerilog (BSV). Because much of what we do involves the business of complex concurrency, lots of moving parts that are difficult to reason about in the whole but easier to grasp a rule-step at a time, BSV has been an obvious choice for us. But there is an important aspect of designing with BSV that we feel is often overlooked by those not familiar with the art, and is perhaps one of the language’s best features:
BSV is the Ultimate Digital Designer’s Assistant!

There’s a saying in our shop with regard to clearly-written BSV codes that goes something like this: “If it compiles, it will work”. With those words, hundreds of verification engineers have just aimed tactical nukes towards Manchester New Hampshire, so please let me explain. There are several reasons why this is the case, I’ll touch on a couple.

First, there are crazy-good type safety and formal guarantees enforced by the compiler. Compared to Verilog, which is awful, and even VHDL, which is less-awful, the BSV compiler does an amazing amount of analysis and type-checking statically at compile time within seconds of your pinky coming down on the enter key. We have spent days debugging issues in conventional RTL that the BSV compiler would have detected in seconds. In 2006, Stuart Sutherland presented a SNUG paper highlighting 57 “Gotchas” with Verilog and SystemVerilog. BSV detects or avoids the vast majority of them. There is simply that much less to get-wrong, and when it is wrong, you are getting the most-excellent feedback from the compiler while the ideas and concepts have barely left your fingertips and are fresh in your mind.


Next, there is this quality; you may have seen it on the masthead of this blog, of “Scalable Atomicity”. In a nutshell, the BSV designer reasons about rules within a module, and methods a module exposes to others. The BSV compiler ruthlessly checks and ensures that the deterministic behavior of the resultant sequential circuit is equivalent to the result of each rule firing one at a time. We look at any rule in our system, and the interesting state-changing actions between the curly-braces and reason about them one-at-a-time. Our reptilian brains can handle that focused task! The BSV compiler then composes a deterministic schedule, devoid of any new state elements, that produces results identical to the one-at-a-time firing of all the rules in the system we have just reasoned about.

Our IP shop designs DMA engines, Packet Processors, Beamformers, and the like. Our clients demand that we “go deep quick”. Meaning we show them something partially-functional almost immediately. Without BSV by our side to check our work at every compile, we would not be able to move nearly as quickly. The edit-compile-debug cycle with BSV is seconds. We go around that loop, including bit- and cycle- true C simulation dozens of times each day. Then we make a bitstream, and it “just works”. If you find yourself complaining that the FPGA vendor tools take a long time and slow your productivity; perhaps you should explore how a digital designer’s assistant such as BSV would change your world.

Wednesday, November 27, 2013

Debian 7.2 and OpenCPI

A client of ours using OpenCPI had selected the Debian flavor of Linux for their work. That's fine. Certainly in the FPGA space, and for that matter, across all processor technologies, we have tried not to dictate any particular Linux. That said, the developers had added support for RedHat, CentOS, and MacOS from the ancient times. More recently Ubuntu was added into the mix. And just yesterday, with this commit, we dealt with Debian. It's interesting seeing the subtle, and not-so-subtle, differences across a dozen different Linux distributions.

Wednesday, October 16, 2013

OCP-IP move to Accellera

We're cautiously optimistic, maybe even excited, about the news that went public yesterday of OCP-IP being rolled into Accellera's warm arms. We have participated in technical issues for both; but very much like that Accellera is a proven conduit to IEEE standardization. Open and accessible standards are a good thing. Time will tell, but out of the gate we are pleased that our man years of investment into the OpenCPI Worker Interface Profiles (WIPs), which often incorporate OCP-IP concepts, can live on. Come see us at FPGA-2014 in Monterey to understand why, years later, we still feel the OCP/AXI choice is like Coke and Pepsi.

Tuesday, October 8, 2013

Ubuntu 13.04 and OpenCPI

We've been using RHEL as our standard Linux for over five years. Initially RHEL5 64b WS, then RHEL6. On the plus side, RHEL is almost always embraced by our CAD and CAE tool vendors as a supported OS. On the minus side, the libraries are as old as the hills and we really weren't feeling the love of sending $200/machine-image/year to Red Hat for this configuration grief. We had noodled in earlier versions of Ubuntu and had some mixed feelings. With Ubuntu 13.04 this past summer, I tried it again and it was just great: everything "just worked". We didn't even have to muck around with GPU drivers. What used to be a half day project of dependency-hell installed in one apt-get. Wow! I don't think of us as OS-bigots; but the understandable contrast between the library hassles we had with something as old as RHEL5 and didn't have with an OS as modern as Ubuntu 13.04 was just too much to overlook. We've started to let our RHEL licenses lapse; just keeping around a few frozen RHEL5s and one maintained RHEL6 so we have them in shop. But unless a client or application needs some other Linux; it's Ubuntu 13.04 in our shop this fall and going forward.

Jim Kulp at Parera did us all a solid by refreshing the OpenCPI mainline to build cleanly to Ubuntu 13.04 as well. Nice! He pushed those changes to the OpenCPI GitHub repo this evening. Thanks Jim!

Sunday, June 30, 2013

Port Mirroring

We have FPGAs carrying on conversations with other FPGAs all the time. Frequently we use layer 2 Ethernet between known devices on a LAN, although we are working our way up to layer 3. Wireshark is a go-to tool to see what is going on and share pcap captures with our colleagues. But how do you see the traffic between two FPGAs on a LAN when they are talking to each other? Port Mirroring!

We had been playing all kinds of games before we found this little gem: the Netgear GS105E. $60 to Amazon, a dorky Windows configuration dance, and we feel we've found just the ticket for our packet capture needs. We set ours up to only mirror the ingress from ports 1 and 2 to port 5. We plug our chatty FPGA board into ports 1 and 2 and anything running Wireshark into port 5. Ta dah. No more sending broadcast packets just to see what is going on!

Tuesday, May 14, 2013

4DSP FMC116

We have used 4DSP products before. We recommend them to clients. We've evaluated several different models and liked the FMC150 so much, we purchased one to keep. We recently had the chance to work with the FMC116, a 16 channel, 125 MSPS FMC module. We were happy, but not surprised, that it took less than a day to bring up in our lab. This is using their supplied user guide; but our own, homegrown BSV codes. I would not want to say that operationalizing any FMC module is easy; but through the work and doc 4DSP put into this quality product, ... it was! Thanks Pierrick and team!!!

4DSP FMC116 On bench

Sunday, April 7, 2013

Vivado 2013.1 / ISE 14.5

Xilinx shipped Vivado 2013.1 last week. If the engineering design and verification community needed any validation that "FPGA CAD Tools are closing ground on their ASIC centric brethren", Vivado is an excellent example. We see so much runway and room to grow here that we've been making proactive investments around nascent aspects of Vivado, such as IP Integrator (IPI). Our efforts in this regard were mentioned in this press release.

While most of our clients are in the heart of design cycles with 28nm silicon; we are also supporting a user base with legacy silicon, particularly Xilinx 6-series, such as the olde-timey, four year old,  ML605 platform. To this end, ISE has bumped from 14.4 to 14.5. We are progressing through our regression tests for OpenCPI and others without drama.

Sunday, March 3, 2013

Component Based Design

We've been working with beta versions of Vivado IP Integrator from Xilinx. Taken at baseline, this is a fine way to compose component assemblies of IP. Out of the box, graphically, it feels like a grown-up System Generator for the masses, not just for DSP.

The choice of an industry standard for component metadata, IP-XACT, helps strengthen the value-proposition. (Disclosure: We are observers on the Accellera Technical Committee).

Perhaps our greatest excitement for this technology is the ability to package our own and our client's IP components. In this manner, they too become first-class citizens in an IP catalog standing alongside the IP catalogs of others.

We would be careful about using phrases like "Correct by Construction" or "Plug and Play". But the fact is that well-defined interfaces with strong guards and type-checking help the situation.

Wednesday, January 23, 2013

Vivado 2012.4 Digilent USB/JTAG with RHEL6 64b

Getting the Digilent bits to work properly with Vivado 2012.4 with RHEL6 64b can be a bit of a dance. Clear webcase reporting to Xilinx is the best way to help drive this issue to ground. Here at AR Auburn, we have refined our procedure for the delta on top of the vanilla 2012.4 install:

Follow the AR42728 . Note that it is now applicable to 14.4/2012.4. The current version of libusb at this writing is 1.0.9, no big deal.

For the next steps, take all the default options...
  • cd to /opt/Xilinx/14.4/ISE_DS/ISE/bin/lin64/digilent
  • cd to /digilent.adept.runtime_2.9.9-x86_64, run the install script with root permission
  • cd to /ftdi.drivers_1.0.4-x86_64, run the install script with root permission
  • cd to /libCseDigilent_2.2.10-x86_64, run the install script without root permission
We have found this gets us user (non-root) access for both impact and xsdk with the KC705, VC707, and ZedBoard. If it doesn't work for you, well, there should be a webcase in your future.

Update 2013-06-11: Be sure to look at the updated Xilinx Answer record AR54382. With Vivado 2013.2 just about out the door, this is where you want to start to get your hardware session straight!

Tuesday, January 8, 2013

Vivado 2012.4

For three weeks now we have been all over Xilinx' release of Vivado 2012.4 . We are excited about this release for several reasons. We have been using 28nm 7-series silicon for sometime with the KC705 and VC707 TDPs; and soon we will increase Zynq's participation in the mix. Although only in beta test at this point, the rapidly maturing functionality of IP Integrator got our attention as well. While running native on RHEL6.3 we've seen zero crashes; and love that we can have any mixture of command-line, scripted and GUI build awesomeness.